магнит, розничная сеть. it

Инженер DevOps (Kafka) /Kafka Engineer

Ищем DevOps / Kafka Engineer в команду потоковых данных, которая развивает и поддерживает платформенное решение на базе Kafka и Kafka Connect. Команда отвечает за отказоустойчивость, доступность и мониторинг Kafka-инфраструктуры, а также а…

магнит, розничная сеть. it · На сервисе с: 28.09.26 15:27

Зарплата не указанаРоссияМоскваУдалёнка

Ищем DevOps / Kafka Engineer в команду потоковых данных, которая развивает и поддерживает платформенное решение на базе Kafka и Kafka Connect. Команда отвечает за отказоустойчивость, доступность и мониторинг Kafka-инфраструктуры, а также автоматизирует управление жизненным циклом кластеров и развивает внутренние инструменты для работы с платформой.

Чем придется заниматься

  • Обеспечивать отказоустойчивость, стабильность и доступность кластеров Kafka и Kafka Connect
  • Развивать и поддерживать платформенное решение на базе Kafka и Kafka Connect
  • Конфигурировать и сопровождать коннекторы MirrorMaker2, Debezium и другие Kafka Connect-коннекторы
  • Заниматься диагностикой и дебагом Kafka, Kafka Connect и связанных инфраструктурных компонентов
  • Автоматизировать развертывание, конфигурирование и масштабирование кластеров с помощью Ansible
  • Развивать внутренние инструменты для автоматизации жизненного цикла Kafka-инфраструктуры
  • Настраивать мониторинг, алертинг и дашборды с использованием VictoriaMetrics, Alertmanager и Grafana
  • Участвовать в миграции данных между кластерами и обеспечивать ее без простоя сервисов
  • Консультировать команды разработки по лучшим практикам работы с Kafka и настройке producers / consumers
Мы ожидаем:
  • Опыт работы DevOps / Infrastructure Engineer от 5 лет, из них не менее 2-3 лет работы с Kafka в production
  • Глубокое понимание архитектуры Kafka и принципов построения отказоустойчивых высоконагруженных кластеров
  • Практический опыт эксплуатации Kafka Connect, MirrorMaker2 или Debezium
  • Уверенные навыки администрирования Linux и диагностики инфраструктурных проблем
  • Опыт автоматизации развертывания и управления инфраструктурой с использованием Ansible
  • Опыт настройки мониторинга, метрик и алертинга, а также работы с VictoriaMetrics, Alertmanager и Grafana
  • Самостоятельность в принятии технических решений и умение консультировать команды разработки по работе с Kafka

Похожие вакансии DevOps

Сайты компаний
T

DataOps Intern

tabbyНа сервисе с: 03.10.26 18:46
Зарплата не указанаСаудовская АравияRiyadhОфис

About the role

About the company

Tabby builds financial products used by millions of users across the GCC. The infrastructure behind them runs at scale, under strict requirements for reliability, cost efficiency and regulatory compliance.

This is not a course and not a shadowing programme. It is an engineering role with real responsibility.

Context

The Data Platform team runs the infrastructure that AI, ML and data workloads at Tabby depend on: compute, orchestration, deployment, observability and cost control across cloud environments. The work sits between classic DevOps and the machine-learning side. The same clusters, pipelines and monitoring that keep a service alive also keep models trained, served and measured.

The internship is designed for strong early-career engineers who are comfortable in Linux and a cloud, and who already use AI tools in their own work rather than reading about them. Interns join the team, work on real production infrastructure under senior review and are expected to meet engineering standards from day one.


Responsibilities

What you will do

This is not a helper or ticket-closing role. Interns work on real production tasks under senior review.
  • Work with the cloud infrastructure (primarily GCP) and the bare-metal fleet that host our data, ML and AI workloads
  • Build and maintain CI/CD pipelines for services and models
  • Run and troubleshoot containerised workloads on Kubernetes
  • Set up and improve monitoring, alerting and logging, and act on what they show
  • Automate repetitive operational work with Python or Bash instead of repeating it
  • Support model training and inference workloads: environments, resources, deployment, cost
  • Investigate incidents in infrastructure and pipelines and help find root causes
  • Improve the reliability and cost efficiency of the platform
  • Work within SAMA regulatory requirements: in Saudi fintech, where data lives and who can reach it is part of the engineering problem, not paperwork someone else handles

What you will actually work with

Not a wish list. This is the stack the team runs today. Nobody is expected to arrive knowing all of it.
  • Data: CDC pipelines, BigQuery, Airflow
  • ML: Airflow, ClearML and similar orchestration and experiment tooling
  • AI: bare-metal GPU servers, vLLM, open-source models served in-house
  • Platform: GCP, Kubernetes, Linux, networking
  • Context: SAMA regulations

Qualifications

Required

  • Solid Linux fundamentals: filesystem, processes, permissions, networking basics, comfortable in the shell
  • Hands-on experience with at least one cloud provider, evidenced by something you actually built or deployed. We run on GCP, so GCP experience is the most directly useful, but AWS or Azure evidence counts: the concepts transfer, and we would rather have someone who has really built something on one cloud than someone who has clicked around ours
  • Understanding of networking: DNS, TCP/IP basics, load balancing, what happens between a request and a service
  • Working knowledge of containers, and enough Kubernetes to deploy and debug a workload
  • Familiarity with monitoring and observability concepts: metrics, logs, alerts and what makes an alert useful
  • Python or Bash sufficient to automate operational tasks
  • Experience with Git and standard development workflows
  • Real, current use of AI tools in your own engineering work: which tools, for what, and an informed view of which models suit which task. We would rather hear an honest comparison than a list of names
  • Structured thinking and attention to correctness
  • Open to constructive feedback
  • English sufficient for documentation and team communication

Strong plus

  • Infrastructure as code (Terraform or similar)
  • Experience running a CI/CD system end to end (GitLab CI, GitHub Actions or similar)
  • An observability stack in practice: Prometheus, Grafana or equivalents
  • Exposure to MLOps tooling: experiment tracking, model registries, feature stores, inference serving
  • Has tried to run an open-source model themselves (on a laptop, a rented GPU, anything) and can explain how LLMs actually work rather than just which API they called
  • Any experience with GPU workloads, or with the cost side of running them
  • Interest in platform design and developer experience

Eligibility

  • Saudi nationals only
  • We welcome both current students and fresh graduates
  • We expect a full-time level of engagement. The programme is not part-time. Students can align time for classes or exams with their mentor in advance, but performance, ownership and involvement are expected at a full-time level

Benefits

  • Six months, starting autumn 2026
  • Paid internship, funded by Tabby
  • Full integration into an engineering team
  • Distributed engineering team across multiple countries
  • A path to a junior role on the platform side afterwards. That is our intent and what we aim for, not a guarantee: it depends on how the internship goes
This internship is intentionally demanding and designed for candidates aiming for fast professional growth in infrastructure and AI platform engineering.

Сайты компаний
T

Senior DevOps Engineer

tabbyНа сервисе с: 03.10.26 18:01
Зарплата не указанаСаудовская АравияRiyadhОфис

About the role

Tabby creates financial freedom in the way people shop, earn and save by reshaping their relationship with money. Over 15 million users choose Tabby to stay in control of their spending and make the most out of their money.

The company’s flagship offering allows shoppers to split their payments online and in-store with no interest or fees. Over 40,000 global brands and small businesses, including Amazon, Noon, IKEA, and SHEIN use Tabby to accelerate growth and gain loyal customers by offering easy and flexible payments online and in stores.

Tabby generates over $10 billion in annual transaction volume for its partner brands and is the highest-rated, most-reviewed, largest, and fastest-growing FinTech in the GCC region.
Tabby launched in 2019 and has since raised +$1 billion in equity and debt funding from global and regional investors, and is now valued at $3.3 billion.

We are seeking a skilled IT professional to join our team in Saudi Arabia. The role involves a variety of responsibilities, including infrastructure maintenance, communication with regulatory authorities, and managing interactions with the Saudi Arabian Monetary Authority (SAMA).

Important note: please submit your CV in English as recruiters do not speak Arabic. Thank you! 

Responsibilities

We are looking for candidates with varying levels of experience, from hands-on administrators to managerial professionals with a technical background.

You'll get to: 
  • Manage and maintain servers and systems (Linux).
  • Attend meetings with regulatory representatives and engage in clear, professional communication. 
  • Perform infrastructure maintenance tasks in Saudi Arabia. 
  • Design scalable and reliable cloud infrastructure
  • Create a cloud platform suitable for running software services
  • Automate infrastructure provisioning
  • Ensure the smooth operation of and participate in emergency response situations for outages in our infrastructure
  • Security hardening (PCI-DSS)
  • Observability and monitoring improvement

Qualifications

• Understanding and hands-on experience in DevOps practices, with hands-on familiarity with Linux (maintenance rather than building automations or pipelines). 
• Experience with VMware, Citrix, or similar platforms is a great advantage.
• Ability to identify regulatory requirements and effectively communicate with Saudi Arabia’s Central Bank.
• Spoken and written Arabic language.

Our tech stack:
Linux, Kubernetes, GCP (GKE, Pub/Sub, Cloud PostgreSQL, Spanner), Datadog, Gitlab, Cloudflare, Elasticsearch, Terragrunt, Istio, Vault, ArgoCD, Helm, Go

Benefits

  • A working environment that gives you autonomy and responsibility from day one.
  • Great learning opportunity to work with world-class Senior DevOps.
  • Participation in the company’s employee stock options program.
  • Competitive salary and other bonuses
  • Health Insurance
We are passionate about creating an inclusive, high-performing workplace that gives people from all backgrounds the support they need to thrive, grow, and meet their goals (whatever they may be).

If this sounds exciting to you, we’d love to hear from you!

Сайты компаний
T

Senior Devops Engineer

taxdomeНа сервисе с: 03.10.26 18:07
Зарплата не указанаНе указана странаЛокация не указанаУдалёнка

About TaxDome

At TaxDome, we’re building the #1 practice management platform for accounting firms in the US and globally. Founded in 2017, we’ve grown into a around 400 fully remote team across 40+ countries, serving tens of thousands of businesses worldwide with millions of end clients.

How we work

You’ll be part of a globally distributed team built on trust, ownership, and self-management.

We focus on outcomes over activity — prioritizing clear ownership, pragmatic decision-making, and accountability for results over rigid processes. Collaboration is central to how we operate: we communicate openly, involve the right people early, and continuously improve how we build products and teams together.

About this role

As the product and infrastructure grow, we are looking for a talented and experienced Senior DevOps Engineer who is comfortable working remotely, demonstrates strong self-organization, and is actively engaged in team collaboration.

It’s a fully remote role, we are hiring across European timezones.

Our tech stack:

  • Ruby on Rails, React.js
  • PostgreSQL
  • AWS (VPC, IAM, EKS, RDS, OpenSearch, CloudWatch, ElastiCache, CloudFront, AWS Secrets Manager)
  • Kubernetes
  • GitLab CI, Helm
  • Autoscaling & Capacity Management: Kubernetes Autoscaler, Karpenter
  • Cloudflare (CDN, DNS, WAF, security and traffic management)
  • Monitoring & Observability: Prometheus Stack, VictoriaMetrics, Datadog, AWS CloudWatch
  • Unix/Linux

What you’ll be responsible for

  • Design, build, and operate reliable production infrastructure for a high-load SaaS product;
  • Own the full lifecycle of environments: development, staging, and production;
  • Develop and evolve Kubernetes-based platform (EKS), including networking, security, scaling, and capacity management;
  • Implement and maintain Infrastructure as Code using Terraform as a single source of truth;
  • Build, maintain, and improve CI/CD pipelines using GitLab CI and Helm;
  • Participate in system architecture design, propose and implement infrastructure and platform improvements;
  • Ensure high availability, performance, and fault tolerance of services;
  • Implement and evolve observability: metrics, logs, traces, alerting, and SLO-driven monitoring;
  • Actively participate in production operations: incident response, root-cause analysis, postmortems, and preventive actions;
  • Work closely with development teams, providing DevOps expertise during design, debugging, and delivery phases;
  • Drive DevOps best practices across teams: automation, standardization, security, and operational excellence.

Examples of interesting tasks 

  • Disaster Recovery — building a ready-to-go recovery process for every service, plus HA failover between availability zones across all services
  • Multi-Region Disaster Recovery
  • Growing our service infrastructure and CI/CD — the company ran on a monolith for a long time; we're now moving to a hybrid architecture
  • Service Mesh
  • Switching incident-management platforms and building the processes around it
  • Moving to KEDA as a single autoscaling layer based on custom metrics
  • Migrating remaining infrastructure workloads to ARM


What you bring

Must-have

  • 6+ years of experience in a DevOps / SRE role;
  • Strong hands-on experience managing production infrastructure in AWS;
  • Confident knowledge of Kubernetes (cluster deployment, configuration, administration);
  • Experience building and maintaining CI/CD pipelines for product teams
  • Solid experience with Infrastructure as Code approaches (Terraform);
  • Experience designing and operating autoscaling and capacity management solutions;
  • Ability to set up and operate monitoring, logging, and alerting systems;
  • Experience with infrastructure and application security, access control, secrets management, and security best practices;
  • Cloud cost optimization and capacity planning (FinOps mindset);
  • Proactive ownership of infrastructure evolution and technical improvements beyond assigned tasks;
  • Strong troubleshooting skills and a high level of ownership.

What we offer

At TaxDome, we aim to create an environment where people can do their best work and grow alongside the company.
  • Competitive compensation, paid in USD, shared on Recruiter screen
  • Performance-based incentive plan
  • Fully remote work with flexible hours
  • 30 paid days off annually, plus sick days as needed
  • Health & well-being support
  • Learning & development budget to support your professional growth
  • English lessons reimbursement
  • Co-working space reimbursement 
  • Company-provided equipment (conditions may vary depending on the role)
  • A high level of autonomy and ownership in your work
  • The opportunity to make a real impact in a fast-growing global SaaS company
  • A collaborative, international team with a strong product mindset
Upon successful completion of the interview process and acceptance of the offer to join, an employment verification check will be conducted as part of pre-boarding — confirming job titles and dates of engagement with 2 of your previous employers. This is a required step for all new joiners.

If you’re excited about this opportunity and believe you could make an impact at TaxDome, we’d love to hear from you.

By applying, you acknowledge that your personal data will be processed in accordance with TaxDome’s Privacy Notice.

HireSeeker собирает вакансии со всех площадок и присылает только релевантные. Бесплатно.