A

Software Engineer – ML Platform (Python)

The ML Platform team at Avride builds the infrastructure that powers large-scale ML training and data processing for autonomous driving. We sit between Cloud Platform and ML engineers, turning low-level compute, storage, and networking pri…

Avride На сервисе с: 03.10.26 15:36

↑ Вакансия с автоподнятием
Зарплата не указанаСШАAustinУдалёнка

About the team

The ML Platform team at Avride builds the infrastructure that powers large-scale ML training and data processing for autonomous driving. We sit between Cloud Platform and ML engineers, turning low-level compute, storage, and networking primitives into an ML platform that teams actually use — scalable orchestration, distributed compute, and production-grade tooling for the full model lifecycle.

About the role

As an ML Platform Engineer at Avride, you'll own critical pieces of the ML stack: workflow orchestration, distributed execution, resource governance, performance.You will shape how ML teams across the company run experiments and train models at scale. You will build the abstractions and services that make training workloads reliable, cost-efficient, and fast, helping ML teams run at scale on Kubernetes with strong reliability and excellent developer experience.

What you will do

  • Build and scale our ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration
  • Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution, and multi-tenant resource governance — scheduling, priorities, quotas, and policy enforcement across GPU, CPU, memory, and IO
  • Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, and removing bottlenecks in storage, network, and resource contention
  • Work directly with ML teams to debug complex workload issues, drive root-cause analysis, and turn recurring problems into platform-level fixes
  • Evaluate, integrate and extend open-source tooling (Argo Workflows, Ray, Kubernetes ecosystem) to meet evolving platform needs

What you will need

  • Strong proficiency in Python or Go; C++ is a plus
  • Track record of designing and building scalable, maintainable systems and services
  • Experience operating production services end-to-end: APIs, reliability practices, observability
  • Deep knowledge of Kubernetes: how scheduling, resource management, controllers, and pod lifecycle actually behave under pressure
  • Solid Linux and systems debugging skills: performance investigation, networking, storage/IO
  • Ability to troubleshoot complex production issues across logs, metrics, and traces and drive them to resolution

Nice to have

  • Experience with Argo Workflows, Ray, MLflow, or comparable distributed ML tooling
  • Hands-on experience building or operating large-scale ML training systems: GPU scheduling, distributed training, training data pipelines
  • Track record of optimizing resource usage and performance in distributed environments

#LI-MS1

 

Candidates are required to be authorized to work in the U.S. The employer is not offering relocation, sponsorship, and remote work options are not available.

Avride is an equal opportunity employer and committed to providing reasonable accommodations to qualified applicants and employees with disabilities to ensure they have equal access to employment opportunities. Avride complies with the Americans with Disabilities Act (ADA), if you need a reasonable accommodation to assist with the application or hiring process, or to perform the essential functions of a job, please email jobs@avride.ai.

Похожие вакансии MLOps инженер

Сайты компаний
S

Machine Learning Lead

social discovery groupНа сервисе с: 03.10.26 17:55↑ Вакансия с автоподнятием
Зарплата не указанаСербия

Social Discovery Group (SDG) is a group of social discovery companies. SDG solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal. SDG’s products redefine the way people interact and connect with one another.


Our portfolio includes social entertainment platforms designed to connect people online across different cultures and regions of the world.


We bring together a team of like-minded people and IT professionals who specialize in creating and developing globally impactful social discovery products. Our international team of digital nomads works remotely from all over the world.


We’re proud to be a two-time “Great Place to Work” winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).


We are looking for Machine Learning Lead.


Your main tasks will be:

  • Own the ML strategy for dialogue systems: decide what to build and in what order, and tie those bets to business metrics (ARPU, retention, chat depth).

  • Lead a team of 3 ML engineers — set the technical bar, distribute work, hire and let go, grow the people you keep.

  • Drive LLM post-training and model adaptation end to end: SFT, LoRA/QLoRA, preference optimization (DPO / ORPO / SimPO / GRPO), dataset construction.

  • Design agent harnesses and tool-using LLM systems: tool calling, structured outputs, routing, retries, memory and context, guardrails.

  • Build the evaluation layer: offline and pairwise evals, LLM-as-a-judge, regression suites that actually correlate with A/B outcomes.

  • Cut dialogue failure modes — loops, contradictions, persona drift, context loss, generic replies — and keep inference efficient on latency and cost per message.

  • Stay hands-on where it matters (roughly 10–20% of your time): prototypes, debugging agent traces, reviewing your team's work.


We expect from you:


  • Technical degree and a real ML engineering background — you have trained and shipped models yourself, not only managed people who do.
  • Experience leading a team of 3–5 ML engineers, including hiring and performance decisions.
  • Expert-level Python, solid understanding of transformer architecture and modern LLM behavior, hands-on with training, fine-tuning and evaluation.
  • Production experience with LLM systems: APIs, streaming, batching, fallbacks, cost/latency trade-offs, observability over traces and transcripts.
  • Ability to read fresh research and turn it into a prototype, an eval and a shipped change with measurable impact.
  • Fluent Russian, ready to work in CET (±2) hours.

Nice to have: experience beyond text (computer vision, image generation, multimodal), long-running conversations and character consistency, vLLM / TGI / SGLang, DeepSpeed / FSDP / Accelerate, quantization, safety classifiers.


What do we offer:


  • REMOTE OPPORTUNITY to work full-time;
  • The initial pay level or pay range for this role will be shared with candidates during the recruitment process and before the commencement of employment;
  • Vacation 28 calendar days per year;
  • 7 wellness days per year (time off) that can be used to deal with household issues, to lie down and recover without taking sick leave;
  • Bonuses up to $5000 for recommending successful applicants for positions in the company;
  • 50% payment for professional training, international conferences, and meetings;
  • Corporate discount for English lessons;
  • ​Health benefits. According to the paychecks, if you are not eligible for corporate medical insurance, the company will compensate you with up to $ 1,000 gross per year per employee. This can be spent on self-purchase of health insurance or on doctor’s fees for yourself and close relatives (spouse, children);
  • ​Workplace organization. The company provides all employees with an equipped workplace and all the necessary equipment (table, armchair, wifi, etc.) in our offices or co-working locations. In the other locations, the company provides reimbursement of workplace costs up to $ 1000 gross once every 3 years, according to the paychecks. This money can be spent on the rent of the co-working room, on equipping the working place at home (desk, chair, Internet, etc.) during those 3 years;
  • Internal gamified gratitude system: receive bonuses from colleagues and exchange them for our merchandise, team building activities, massage certificates, etc. 

Sounds good? Join us now!

Сайты компаний
ozon

Ведущий разработчик ML, Рекомендации

ozonНа сервисе с: 06.07.26 22:57↑ Вакансия с автоподнятием
Зарплата не указанаРоссияМосква

Привет! Мы команда Сервисы формирования лент.
Мы строим платформу, которая превращает терабайты данных в персонализированные рекомендации для миллионов пользователей Ozon. Наши алгоритмы решают, что покупатели увидят в следующий момент, и в этом нам помогают сложные ML-модели, работающие в реальном времени. Вместе с командой Data Science мы постоянно улучшаем платформу, делая рекомендации точнее и релевантнее.

Сейчас мы ищем ведущего ML-инженера, который возьмёт на себя проекты ключевого сервиса инференса моделей и расчёта признаков. Вы будете отвечать за техническую экспертизу, глубоко вникать в архитектуру и разработку и помогать команде справляться с нетривиальными задачами.

Наш стек

Go (основные сервисы), Python (ML-часть), Cassandra, ScyllaDB, Redis, NVIDIA Triton, ONNX, OpenVINO, Kubernetes, Hadoop, Airflow.


Вы будете

  • Проектировать, разрабатывать и оптимизировать высоконагруженные ML-сервисы для рекомендаций, нейросетевого скоринга и обработки стриминговых данных.

  • Интегрировать решения с ML-платформой и инфраструктурой Ozon.

  • Продуктивизировать сложные модели, включая нейросети и аналоги LLM.

  • Улучшать пайплайн публикации моделей и метрик качества рекомендаций.

  • Внедрять мониторинг, обеспечивать отказоустойчивость и низкие задержки.

  • Заниматься развитием коллег в команде, участвовать в код-ревью, делиться экспертизой и помогать команде принимать технические решения.

Нам важно

  • Глубокие знания в ML Engineering и работе с высоконагруженными системами.

  • Опыт вывода в production сложных моделей, включая нейросетевые подходы (LLM, глубокие рекомендательные системы и т.д.).

  • Умение проектировать распределённые системы и оптимизировать их производительность.

  • Способность вести экспертные обсуждения, разъяснять сложные решения и вдохновлять команду.


Будет плюсом

  • Опыт работы с LLM, рекомендательными системами или скорингом.

  • Практика внедрения стриминговой обработки данных.

  • Понимание принципов персонализации в e-commerce или маркетплейсах.

Наши сервисы написаны на Go, но мы открыты для специалистов с опытом на других языках (Python, Java, С++, C#, Rust и т.д.), готовых погрузиться в наш стек. Главное – экспертность в ML Engineering и желание решать сложные задачи в команде!

Соц.сети
Н

Python Engineer (ML/Infrastructure)

Неизвестный работодательНа сервисе с: 04.10.26 21:04
Зарплата не указанаСШАУдалёнка

Hiring - Remote role, US-based.

Python Engineer (ML/Infrastructure)

We're building the platform layer behind Cloud AI Platform: the tooling, services, and developer experience that AI-driven systems run on. We need a Python engineer who's comfortable on both sides - ML and infrastructure.

What you'll be doing:

- Partnering with internal product teams, understanding their AI/ML use cases, and turning requirements into technical solutions

- Prototyping fast, validating with customers, then hardening what works for production

- Building services, integrations, workflows, and developer tooling on top of Cloud AI Platform

- Spotting recurring customer needs and turning them into reusable platform capabilities

- Working with platform teams to improve APIs, SDKs, docs, and the developer experience

What we're looking for:

- Solid Python skills. Go or Scala is a plus

- Experience designing, building, and maintaining ML infrastructure and deployment pipelines

- Docker and Kubernetes

- Cloud platforms: AWS, Azure, or GCP

- IaC and CI/CD: Terraform, CloudFormation, Jenkins, GitLab CI, GitHub Actions

- Orchestration and streaming: Airflow, Prefect, Dagster, Kafka, Kinesis

- Monitoring and observability: Prometheus, Grafana, ELK stack, for both model performance and system health

- Strong software engineering fundamentals and DevOps practices

- Git and collaborative development workflows

- 3+ years in MLOps, DevOps, or related infrastructure roles

- BS or MS in Computer Science, Software Engineering, Machine Learning, or an equivalent degree with applicable experience

Nice to have:

- ML frameworks: TensorFlow, PyTorch, MLflow, Kubeflow

- Security best practices for ML systems and data governance

- Model versioning, experiment tracking, feature stores: MLflow, Weights & Biases, Feast

- Automated testing for ML systems, including data validation and model testing

You'll be the bridge between product, platform, and customer teams - defining best practices, documenting patterns, and driving alignment across groups. You'll work in a fast-evolving environment where engineering fundamentals, applied ML awareness, and customer empathy all matter.

Fully remote in the US. If this sounds like your kind of work, DM me.

HireSeeker собирает вакансии со всех площадок и присылает только релевантные. Бесплатно.