📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

This job is no longer available

This job expired on 26/07/2026. It no longer accepts applications.

Senior Site Reliability Engineer – Platform Squad

Flip

Senior 🇬🇧 English
Kubernetes Azure Go Python Pulumi Infrastructure as Code Observability Loki Grafana Tempo Mimir Prometheus ELK SLI SLO Error budgets

Job description

About the role

As a Senior Site Reliability Engineer in Flip’s Platform Squad, you will own critical reliability domains end‑to‑end, shaping the technical direction of the platform. You will lead architectural decisions, mentor teammates, and continuously raise the reliability bar for a high‑throughput, globally‑scaled SaaS product.

Key responsibilities

  • Co‑own the architecture of Azure cloud infrastructure and Kubernetes clusters, ensuring high throughput and availability.
  • Define and implement a resilience strategy covering global scaling, zero‑downtime deployments, rollbacks, and disaster recovery.
  • Evolve the observability stack (Loki, Grafana, Tempo, Mimir, Prometheus, ELK) into a trusted foundation for engineers.
  • Improve the Infrastructure‑as‑Code platform with Pulumi, reducing toil and enabling self‑service infrastructure.
  • Lead major platform incidents, conduct blameless post‑mortems, and drive systemic improvements.
  • Mentor squad members, run RFCs and design reviews, and help engineers grow into stronger SREs.
  • Partner with the squad to shape the platform roadmap and future direction.

Required profile

  • 5+ years of hands‑on experience as an SRE, Platform Engineer, DevOps Engineer, or similar role.
  • Proven track record building and operating high‑throughput, highly available production systems.
  • Deep production‑level experience with Kubernetes on any hyperscaler (Azure preferred).
  • Strong expertise in modern observability stacks and a clear understanding of SLIs, SLOs, and error budgets.
  • Solid software development skills in Go (preferred) or Python.
  • Hands‑on experience with Infrastructure as Code, preferably Pulumi.

Required skills

  • Kubernetes
  • Azure cloud
  • Go
  • Python
  • Pulumi (IaC)
  • Observability tools: Loki, Grafana, Tempo, Mimir, Prometheus, ELK
  • SLI/SLO management and error budgeting
  • High‑throughput system design
  • Incident management and post‑mortem analysis

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Flip.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.
💬 Chat with us on Telegram Chat on WhatsApp

Published 4 months ago

30 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Flip