5-8+ years of hands-on experience in Machine Learning Engineering, specifically focused on building and scaling MLOps infrastructure and productionizing ML systems.Production Infrastructure: Proven expertise in deploying low-latency, high-throughput ML inference services (using FastAPI, TorchServe, Triton Inference Server, or Ray Serve) across both classical lightweight and heavy-width ML models (PyTorch/TensorFlow). Strong preference for AWS (EKS, EC2, SageMaker) / Snowflake and Open Source ecosystems over GCP/Azure.MLOps & Pipelines: Deep experience building automated, continuous model retraining pipelines to handle concept drift (ranging from daily to weekly cycles). You have orchestrated decoupled, multi-model AI architectures using tools like Airflow, Kubeflow, or Metaflow, and possess strong expertise in model registry and tracking tools like MLflow or Weights & Biases.Feature Stores: Hands-on experience evaluating, building, or extensively leveraging online (Redis, DynamoDB) and offline (Snowflake, S3) Feature Stores in a production environment. Familiarity with frameworks like Feast or custom dbt-based pipelines is highly valued.Strategic Builder Mindset: You are an analytical builder who thinks long-term. You can successfully evaluate TCO for bespoke internal systems versus enterprise tools, anticipate technical liabilities, and design robust architectures that handle unpredictable peak traffic surges.Collaboration & Engineering Hygiene: Strong cross-functional communication skills. You excel at translating complex ML prototypes into highly scalable production code backed by strict version control, rigorous testing, and CI/CD best practices, seamlessly connecting data science innovation with backend engineering execution.Nice-to-Haves:
Deadline:22 Oct 2026
