Machine Learning Engineer

Scale.jobs
San Francisco, CA
F

Machine Learning Engineer

Uber
San Francisco, CA
F

Machine Learning Engineer

10a Labs
San Francisco, CA
F

Machine Learning Engineer

Mariana Minerals
San Francisco, CA
F

Machine Learning Engineer

Strava
San Francisco, CA
F

Machine Learning Engineer

Docusign
San Francisco, CA
F

Machine Learning Engineer

Blank Bio
San Francisco, CA
F

Machine Learning Engineer

Mercor
San Francisco, CA
F

Machine Learning Engineer

Latent
San Francisco, CA
F

Machine Learning Engineer

Blank Bio (YC S25)
San Francisco, CA
F

Machine Learning Engineer

Bland
San Francisco, CA
F

Machine Learning Engineer

HealthLeap AI
San Francisco, CA
F

Machine Learning Engineer

krea.ai
San Francisco, CA
F

Machine Learning Engineer

Scribd, Inc.
San Francisco, CA
F

Machine Learning Engineer

TwelveLabs
San Francisco, CA
F
Scale.jobs company logo

Machine Learning Engineer

Scale.jobs

San Francisco, CA

Full-time

Engineering, Science / R&D / Research

About The Role

The role drives the development and scaling of core machine learning services, bridging the gap between experimental prototyping and highly available production systems. The engineer will collaborate closely with product and data platform teams to embed predictive intelligence and generative capabilities into customer-facing software.

The focus spans across traditional ML pipelines, deep learning paradigms, and modern LLM application patterns. Operating in a remote-first, high-autonomy culture, the engineer will make critical architecture decisions regarding model training pipelines, real-time inference optimization, and MLOps tooling.

Key Responsibilities

  • Build, optimize, and maintain scalable machine learning pipelines for model training, validation, and batch or streaming inference.
  • Develop and deploy retrieval-augmented generation (RAG) applications, agentic workflows, and fine-tuning scripts utilizing state-of-the-art LLMs.
  • Implement robust evaluation frameworks and unit testing suites for ML models to detect regression, hallucinations, and performance degradation.
  • Collaborate with backend engineers to expose model inference endpoints via high-throughput, low-latency APIs built with FastAPI or gRPC.
  • Establish automated MLOps infrastructure for model monitoring, data drift detection, and continuous integration/continuous deployment (CI/CD) of ML assets.
  • Optimize inference latency and GPU utilization through quantization, pruning, and model compilation libraries like TensorRT or vLLM.

What We Are Looking For

  • 3-6 years of experience as a Machine Learning Engineer or Software Engineer working on production-grade AI systems.
  • Deep proficiency in Python and solid experience with ML frameworks such as PyTorch, scikit-learn, and Hugging Face Transformers.
  • Hands-on experience with vector search databases (e.g., Pinecone, Qdrant, Milvus, or pgvector) and modern orchestration tools like LangChain or LlamaIndex.
  • Solid understanding of relational and non-relational databases, including experience building feature pipelines in SQL, pandas, or PySpark.
  • Familiarity with containerization (Docker, Kubernetes) and cloud-based ML orchestration platforms like AWS SageMaker, GCP Vertex AI, or Run:ai.
  • Bonus: Experience with Triton Inference Server, Kubernetes-native ML tools (Kubeflow, KServe), or contributions to open-source ML/LLM repositories.

About the company

Company websiteTechnology, Information and Internet

Scale.jobs is a human-powered job application service for serious job seekers in the US market, built to help them move faster, apply smarter, and land sooner.

We apply to jobs for you manually, at scale, and with precision. Our job concierge team manages the full application process on your behalf, so your time goes toward interview prep, not form-filling.

Scale.jobs is built for Early and senior professionals, H-1B holders, OPT students, new graduates, and basically everyone who wants to save their time for upskilling, networking, learning, and competing in a difficult US job market.

Every client gets a dedicated job hunting team that applies to hundreds of targeted roles, matched to their experience, visa status, and career goals.

We also offer resume optimization, LinkedIn profile optimization, and career strategy consulting for job seekers who want every part of their search working together.

The results speak for themselves.

93% of Scale.jobs clients land a job within 3 months. The average job search goes from 5 months to 1–3.
70% of clients who request a refund because they found a job before their credits ran out. That's how our refund policy works - get back everything you didn't use, even after you get hired.

On Trustpilot, we're rated 4.6 across 202 reviews.

Our clients have come from Stanford, Carnegie Mellon, NYU, Lyft, Instacart, and Michigan and the one thing they had in common was not wanting to apply alone anymore.