детский мир. дм-тех

Data Engineer (Hadoop / Spark)

«Детмир Тех» развивает цифровую экосистему группы компаний «Детский мир». Мы создаём продукты для миллионов клиентов и сервисы, которые помогают бизнесу принимать решения на основе данных.

детский мир. дм-тех · На сервисе с: 29.09.26 17:41

Зарплата не указанаРоссияЕкатеринбургМоскваУдалёнка
Локации (удалёнка):ЕкатеринбургМосква

Удалённая работа — географически не привязана. Щёлкни по любой из локаций, чтобы открыть оригинальную карточку.

«Детмир Тех» развивает цифровую экосистему группы компаний «Детский мир». Мы создаём продукты для миллионов клиентов и сервисы, которые помогают бизнесу принимать решения на основе данных.

Ищем Data Engineer в команду Big Data Platform. Платформа построена на Hadoop 3 и Apache Spark и является основным источником данных для всей группы компаний.

Главная задача — развивать надёжную и производительную платформу обработки больших объёмов данных: проектировать пайплайны, оптимизировать Spark-приложения, развивать Data Lake и совершенствовать архитектуру Hadoop-кластера.

Чем предстоит заниматься:

  • Разрабатывать и оптимизировать batch-процессы на Apache Spark.
  • Проектировать ETL/ELT-пайплайны для обработки больших объёмов данных.
  • Создавать и поддерживать витрины данных в Hadoop.
  • Оптимизировать Spark jobs: partitioning, shuffle, data skew, использование памяти и ресурсов кластера.
  • Развивать архитектуру Data Lake и подходы к организации и хранению данных.
  • Организовывать загрузку и передачу данных через Kafka и S3 API (MinIO).
  • Автоматизировать процессы с помощью Apache Airflow и Apache NiFi.
  • Поддерживать трансформации данных в dbt Core.
  • Развивать и поддерживать витрины данных в ClickHouse.
  • Оптимизировать запросы и процессы загрузки данных в ClickHouse, разбираться с проблемами производительности.
  • Поддерживать аналитическую платформу на Apache Superset, решать возникающие технические проблемы.
  • Внедрять проверки качества данных и мониторинг пайплайнов.
  • Исследовать причины сбоев и деградации производительности.
  • Помогать аналитикам и дата-сайентистам эффективно работать с платформой.
  • Участвовать в развитии внутренних инструментов и стандартов Data Engineering.
  • Участвовать в развёртывании и эксплуатации компонентов платформы в Kubernetes.

Что для нас важно:

  • От трёх лет опыта работы с Hadoop и Apache Spark.
  • Знание основных компонентов Hadoop и принципов работы HDFS и Spark.
  • Глубокое понимание архитектуры Spark и принципов распределённых вычислений.
  • Опыт разработки и оптимизации Spark batch jobs.
  • Уверенное владение SQL.
  • Опыт проектирования DWH или Data Lake.
  • Опыт построения ETL/ELT-процессов.
  • Практический опыт работы с Apache Airflow.
  • Знание форматов хранения данных.
  • Понимание принципов партиционирования и организации данных в Data Lake.
  • Умение диагностировать проблемы с shuffle, partitioning, data skew, памятью и использованием ресурсов кластера.
  • Умение анализировать производительность распределённых приложений и находить узкие места.
  • Умение писать чистый, тестируемый и поддерживаемый код.
  • Готовность принимать технические решения и отвечать за результат.

Будет преимуществом:

  • Опыт разработки на Scala.
  • Знание Kafka и принципов потоковой обработки данных.
  • Опыт работы с ClickHouse и оптимизации запросов.
  • Опыт работы с Apache Superset.
  • Знакомство с Apache NiFi и dbt Core.
  • Опыт эксплуатации приложений в Kubernetes.
  • Знание Docker, Helm и GitLab CI/CD.
  • Опыт работы с DataHub или другими Data Catalog решениями.
  • Опыт построения мониторинга data-платформ и пайплайнов.
  • Знакомство с VictoriaMetrics, Grafana, Hue или Jupyter.
  • Опыт работы с большими Hadoop-кластерами и высокими объёмами данных.

Наш стек:

Хранение и обработка: Hadoop 3, HDFS, Apache Spark

Доставка и обмен данными: Kafka, S3

Оркестрация и трансформации: Apache Airflow, Apache NiFi, dbt Core

Аналитические хранилища: ClickHouse

Ad-hoc аналитика: Hue, Jupyter

Каталог данных: DataHub

BI и визуализация: Apache Superset, Grafana

Мониторинг: VictoriaMetrics, Grafana

Инфраструктура: Kubernetes, Docker, Helm

CI/CD: GitLab CI/CD

Мы предлагаем:

  • Официальное оформление в соответствии с ТК РФ в IT-аккредитованную компанию;
  • Совокупный доход: оклад + годовой бонус;
  • График работы: 5/2 (гибридный формат работы);
  • Вакансия открыта для кандидатов, проживающих в РФ, удалённый формат работы из других стран не рассматривается;
  • Техника для работы;
  • Полис ДМС;
  • Программа Best Benefits;
  • Внутреннее обучение;
  • Спортивные и развлекательные мероприятия, участие в корпоративной жизни компании.

Присоединяйтесь к нам и станьте частью истории бренда с историей!

Похожие вакансии Data Engineer

Сайты компаний
V

Database Developer

veeamНа сервисе с: 03.10.26 21:52
Зарплата не указанаПольшаWarsawУдалёнка

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

We are currently seeking a Database Developer to join us in working on our products.

What You’ll Do

  • Development and support of database structures and objects, writing queries and stored procedures (PostgreSQL, MS SQL)
  • Optimization of the existing database structure and code
  • Modernization and creation of new reports, widgets, and dashboards for Reporting products (a new one and Veeam ONE)
  • Migrating and reworking the current business logic to a new, modern architecture in the current product and adapting it for the new product
  • Working closely with teams of analysts and QA

Technologies we work with:  

  • PostgreSQL, pgTAP or analogs, TFS + Git, MS SQL, T-SQL

What You’ll Bring

  • 3+ years of experience with commercial products
  • Hands-on experience with PostgreSQL
  • Strong knowledge of database theory
  • Ability to write complex SQL scripts, functions, and stored procedures
  • Comfortable reading, understanding, and maintaining existing code
  • Flexibility and willingness to learn
  • Ownership mindset: reliable, detail-oriented, disciplined, and able to deliver efficiently

What You’ll Get 

  • 26 paid days off annually, plus 4 extra global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares
  • Paid parental, maternity, and paternity leave
  • Fully covered family medical plan, dental, rehab, and vaccinations
  • Life, critical illness, and disability insurance
  • Employer pension contribution via PPK
  • Monthly Edenred allowance of 450 PLN for meals
  • MultiSport card fully covered by Veeam, giving access to sports facilities nationwide
  • Up to 12 free therapy sessions annually, plus legal and financial advice
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops and learning events like our annual Global Day of Learning

Please note: If the applicant is permanently present outside of Poland, Veeam reserves the right to refuse to consider the application for a job. Remote job is only possible in case the employee is located in Poland.

#LI-VE1

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

Сайты компаний
V

Director, Graph Databases

veeamНа сервисе с: 03.10.26 21:43
Зарплата не указанаСШАКанадаSan Jose

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

You’ll lead two core systems inside Veeam Data Command Center: the Knowledge Graph and the hyperscale data lake integrations. Together, they help customers understand where sensitive data lives, who can access it, how it moves, and whether AI models trained on it can be trusted. 

You’ll lead multiple teams building a searchable, security-aware graph that works at enterprise scale. This role is for a hands-on technical leader who can set clear direction, grow strong teams, and deliver reliable systems—while building an AI-first engineering culture with high standards for quality and security. 

What You’ll Do

  • Set the technical vision and end-to-end architecture for the Knowledge Graph, including the data model, storage engine, and query layer at very large scale 
  • Guide the evolution of the graph schema for data sources, identities, access, classifications, and lineage (property graph and/or RDF) using Amazon Neptune and/or Neo4j 
  • Own the strategy for hyperscale lake and lakehouse integrations, including connectors and scanning engines that ingest metadata and lineage from Delta Lake, Iceberg, Parquet/Avro, and platforms like Azure Data Lake, AWS S3/Glue, and BigQuery without disrupting production 
  • Drive performance and reliability, including standards for indexing, partitioning, and query planning, and tuning traversals and queries (Gremlin, Cypher/openCypher, SPARQL) 
  • Build and scale an AI-first engineering approach where teams use tools like Claude Code, Cursor, and Copilot responsibly, with guardrails for security, maintainability, and code quality 
  • Invest in reusable engineering building blocks (including “Claude skills” and agent workflows) that make teams faster and more consistent 
  • Own delivery outcomes: roadmap execution, operational readiness, incident learning, and cross-team alignment 
  • Hire, coach, and develop leaders, including engineering managers and senior/staff engineers, with clear expectations and growth paths 

What You’ll Bring

  • 10+ years of software engineering experience in data infrastructure, graph systems, or distributed data platforms 
  • 4+ years of engineering leadership experience, including leading through managers and scaling multiple teams 
  • Strong production experience with Amazon Neptune and/or Neo4j, including scaling, operations, and trade-offs (property graph vs. RDF) 
  • Proven ability to lead graph modeling for complex domains, including lineage and permissions at enterprise scale 
  • Deep knowledge of Gremlin, Cypher/openCypher, and/or SPARQL, including performance tuning and query design best practices 
  • Experience with data lakes/lakehouses (Delta Lake, Iceberg, Parquet) across major cloud platforms (Azure Data Lake, AWS S3/Glue, BigQuery) 
  • Experience designing and operating distributed systems using tools like Spark, Flink, or Presto/Trino, with strong judgement on scalability and cost 
  • Strong backend background in Go and/or Python, with the ability to review designs, guide decisions, and unblock teams 
  • Practical experience using AI-assisted development tools and the ability to set standards that keep AI-assisted code secure and high quality 

Bonus Skills

  • Experience operating graph systems at massive scale 
  • Background in data security, access governance, and policy controls 
  • Experience with AI/ML governance tools and practices (e.g., MLflow, Databricks Mosaic AI) 
  • Experience building custom agents, MCP-based workflows, or reusable engineering automation 
  • Infrastructure-as-Code experience (e.g., Terraform or Pulumi) 
  • Contributions to graph standards or communities (GQL, openCypher, SPARQL) 

 

What you'll get

  • Unlimited paid time off, 12 paid holidays including 4 global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
  • Medical, dental, and vision coverage starting on your first day
  • Mental health support, therapy sessions, and digital wellness tools via our Employee Assistance Program
  • 401(k) retirement plan with company matching contributions
  • Fertility, adoption, and surrogacy support through Maven, plus paid volunteer time
  • AirVet: 24/7 virtual veterinary care at no cost
  • Legal services, identity protection, and supplemental health insurance options
  • Tax-advantaged spending accounts for healthcare, dependent care, and commuting
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning

Pay Transparency

Veeam is committed to pay transparency and equitable compensation. For this role, the compensation range below reflects the expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus. For roles with a commission plan, the compensation range represents On Target Earnings (OTE), which includes base salary plus variable commission. When determining compensation, Veeam takes into consideration factors such as experience, education, skills, and geographic zone. Offers are typically made below the midpoint of the range.

In addition to compensation, Veeam provides a comprehensive benefits package, including health coverage, retirement plans, and unlimited time off.

Compensation Range (TTC / OTE)
$382,560—$710,400 USD

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

Сайты компаний
V

Senior Engineer - Hyperscale Analytics

veeamНа сервисе с: 03.10.26 21:30
Зарплата не указанаСШАКанадаSan JoseГибрид

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

We are seeking exceptional Senior Engineer- Hyperscale Analytics to design and scale next-generation data processing and analytics platforms that power OLTP, OLAP, and large-scale distributed data systems. You will build and optimize pipelines and services that handle billions of records daily, enabling real-time transactions, analytical insights, and AI-driven decisioning.

What You’ll Do

  • Transactional & Analytical Systems: Design and implement highly scalable OLTP systems for real-time workloads and OLAP systems for complex analytical queries on massive datasets
  • Distributed Processing: Build, optimize, and maintain large-scale batch and streaming pipelines using frameworks such as Apache Spark, Flink, Presto/Trino, or Kafka Streams
  • System Performance & Scale: Optimize systems for low-latency queries, high-throughput ingestion, and interactive analytics, ensuring seamless performance as data volumes scale to petabytes
  • Data Infrastructure: Develop and integrate with modern storage and processing systems (e.g., Snowflake, BigQuery, Redshift, Cassandra, HDFS, Delta Lake, Iceberg) to support hybrid analytical/transactional workloads
  • Reliability & Observability: Ensure high availability, reliability, and monitoring across large compute and storage clusters with automated failover and recovery
  • Collaboration: Partner with data scientists, ML engineers, and product teams to build unified, secure, and cost-efficient data platforms

What You’ll Bring

  • 6+ years of professional software engineering experience, with a significant portion focused on data infrastructure, distributed systems, or large-scale analytics platforms
  • Deep experience with distributed data processing frameworks (e.g., Spark, Flink, or similar)
  • Strong background in cloud-scale data architecture (AWS, Azure, or GCP) — data lakes, warehouses, streaming platforms
  • Proficiency in one or more of: Python, Java, Scala, or Go
  • Experience with modern data warehouse/lakehouse technologies (e.g., Snowflake, Databricks, BigQuery, Redshift)
  • Solid understanding of data modeling, ETL/ELT design, and pipeline orchestration (e.g., Airflow, dbt)
  • Track record of designing for scale, reliability, and cost efficiency in production systems

Bonus Skills

  • Experience with HTAP (Hybrid Transactional/Analytical Processing) systems or real-time analytics platforms
  • Familiarity with data lakehouse architectures and formats like Parquet, ORC, Delta, Iceberg, Hudi
  • Knowledge of containerized deployments (Docker, Kubernetes) and cloud-native data architectures (AWS Redshift, GCP BigQuery, Azure Synapse)
  • Background in query engine development or contributing to open-source OLAP/OLTP frameworks

 

What you'll get

  • Unlimited paid time off, 12 paid holidays including 4 global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
  • Medical, dental, and vision coverage starting on your first day
  • Mental health support, therapy sessions, and digital wellness tools via our Employee Assistance Program
  • 401(k) retirement plan with company matching contributions
  • Fertility, adoption, and surrogacy support through Maven, plus paid volunteer time
  • AirVet: 24/7 virtual veterinary care at no cost
  • Legal services, identity protection, and supplemental health insurance options
  • Tax-advantaged spending accounts for healthcare, dependent care, and commuting
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning

Pay Transparency

Veeam is committed to pay transparency and equitable compensation. For this role, the compensation range below reflects the expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus. For roles with a commission plan, the compensation range represents On Target Earnings (OTE), which includes base salary plus variable commission. When determining compensation, Veeam takes into consideration factors such as experience, education, skills, and geographic zone. Offers are typically made below the midpoint of the range.

In addition to compensation, Veeam provides a comprehensive benefits package, including health coverage, retirement plans, and unlimited time off.

Compensation Range (TTC / OTE)
$234,840—$436,080 USD

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

HireSeeker собирает вакансии со всех площадок и присылает только релевантные. Бесплатно.