🔥 Senior DevOps / Site Reliability Engineer - Healthcare / Voice / Observability
We are currently looking for a Middle+ / Senior DevOps / Site Reliability Engineer to join an international healthcare project focused on workflow and voice services for US health-system use.
The main goal of the role is to make production services observable, reliable, recoverable, and safe to operate.
📍 Locations: Poland, EU countries, Georgia, Uzbekistan, Kazakhstan
💼 Seniority: Middle+ / Senior
🇬🇧 Language: English
About the Project
The platform supports medical voice services and workflow automation for healthcare systems in the US.
The role focuses on production reliability: monitoring, incident response, SLOs, deployment safety, recovery, and uptime evidence.
❗️ MUST HAVE
— Strong production experience in SRE / DevOps / Backend Operations
— Hands-on experience with monitoring, logging, alerting, and tracing
— Real experience with Incident Response / Incident Management
— Clear responsibility for uptime, SLA / SLO / service reliability
— Experience defining or working with SLOs and production dashboards
— Strong cloud infrastructure background
— Experience with containers and deployment automation
— Experience with Infrastructure as Code (IaC)
— Strong CI/CD knowledge
— Python or comparable scripting skills for operational troubleshooting and automation
— Experience diagnosing and resolving production failures
— Understanding of backup, recovery, and capacity testing
— Ability to communicate clearly during incidents and provide concise written updates
Important: the CV should clearly show that you were not only configuring infrastructure, but were directly involved in incident response and accountable for uptime / SLA / SLO.
Responsibilities
— Define service-level objectives for production services
— Build and maintain useful reliability dashboards
— Instrument metrics, logs, and traces
— Monitor workflow, voice, and third-party service dependencies
— Lead incident response during agreed coverage hours
— Write clear post-incident reviews
— Improve deployment safety with rollback and staged releases
— Maintain change records and release processes
— Test capacity, backups, and disaster recovery
— Coordinate handover and escalation with engineering teams
— Provide accurate uptime and reliability evidence
— Support continuous improvement of platform resilience
Strong Advantage
— Healthcare or another regulated production environment
— WebRTC / real-time media / voice infrastructure
— LLM-backed services
— Experience handling third-party provider outages
— Grafana / Prometheus / Datadog or similar observability tools
— Experience with customer-facing incident communication
⏰ Working Hours - Important
The team works with US-based stakeholders, and calls are expected during Arizona time (MST, UTC-7).
Candidates should be comfortable with the possibility of regular communication and incident-related coverage during US / Arizona business hours.
💚 What We Offer
— Small-company feel within a fast-growing international environment
— Friendly, collaborative, and mission-driven team
— 25 calendar days of vacation + 5 additional paid sick days
— Medical insurance
— Corporate English courses
— Corporate events and team-building activities
— Support with professional certifications
— Reimbursement for professional courses and training
— Long-term international projects
— Opportunities for professional and technical growth