ID Quantique is hiring a Principal SRE Engineer to lead reliability strategy for its global platform. This hybrid role requires deep expertise in AWS/GCP, observability, and incident management to ensure resilient, secure services.
The Platform Engineering team at ID Quantique is responsible for building, securing, and operating scalable infrastructure that supports cloud-managed SaaS products alongside on-premises components. This principal-level position provides technical leadership to define the standards and operating model for highly available services, requiring a blend of strategic oversight and hands-on involvement with production technologies.
Responsibilities
- Establish and drive the reliability strategy for Tier 1 and Tier 2 services, covering service-level objectives, error budgets, observability, resilience, disaster recovery, and cloud security.
- Design and manage observability platforms, lead chaos engineering and disaster recovery efforts, and enhance the stability of core systems like PostgreSQL, Redis/Valkey, Kafka, and OpenSearch.
- Implement AI-driven operations, including predictive monitoring, automated remediation, and self-healing mechanisms.
- Direct major incident response, on-call rotations, and global escalations while mentoring team members and setting engineering standards.
- Collaborate with Architecture, DevSecOps, Cloud Operations, and Product Development to create a shared reliability roadmap and enforce governance.
- Oversee production reliability, backup strategies, and platform governance across all critical services.
Requirements
- Over 15 years of production engineering experience, including recent hands-on management of large-scale, fault-tolerant systems on AWS or GCP.
- Proven track record of leading resilience initiatives, validating disaster recovery procedures, and managing high-severity production incidents.
- Demonstrated technical leadership through architecture reviews, governance, mentoring, and the establishment of engineering standards adopted across multiple teams.
- Strong command of service-level objectives, error budgets, and observability practices.
Nice to have
- Experience with cloud security posture management, workload protection, policy-as-code, and secure-by-default infrastructure.
- Hands-on expertise in AI traffic management, LLM gateways, capacity planning, resource rightsizing, and FinOps.
- Deep knowledge of networking, routing, load balancing, and distributed systems, with the ability to integrate security and reliability into scalable architectures.
- Familiarity with autonomous operations, predictive alerting, and self-healing platforms.
What the company offers
- Competitive compensation, benefits package, and equity plan.
- Flexible work models to balance professional and personal life.
- Ongoing training and development opportunities.
- Career development plans within the IonQ group’s companies.
- A dynamic, close-knit environment with global expert collaboration.
- A performance-driven culture based on trust, teamwork, and shared ambition.
- Additional benefits including pension, health insurance, public transport pass, meal vouchers, gym membership, parental leave, home office allowance, learning budget, and relocation support.
About the company
ID Quantique is recognized as one of the fastest-growing and most exciting quantum companies globally. The organization fosters a performance-driven culture built on trust, teamwork, and shared ambition, encouraging collaboration among a global team of experts.
- Fastest-growing and most exciting quantum company in the world.
- Performance-driven culture built on trust, teamwork, and shared ambition.
- Dynamic, close-knit environment with global expert collaboration.
Quelle: öffentlich zugängliche Karriereseite des Arbeitgebers. Batchly ist nicht der Arbeitgeber und steht nicht notwendigerweise in einem Vertragsverhältnis mit dem Unternehmen.