📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

Senior Site Reliability Engineer - Ceph / On-Prem Cloud Storage

gridscale · Köln

New
Unbefristet Senior 🇬🇧 English
Ceph OpenStack Kubernetes Ansible Terraform FluxCD ArgoCD Linux KVM Go Python Prometheus Grafana VLAN BGP NVMe BlueStore

Job description

About the role

You will help build, operate and industrialise the storage foundation of our on‑premise cloud platform. As a senior member of a small, experienced team you will own the Ceph clusters that power block, object and file storage for customer workloads, influencing architecture, hardware lifecycle, automation and technical direction.

Key responsibilities

  • Design, operate and evolve Ceph clusters for block, object/S3 and file storage, covering CRUSH topology, pool design, erasure coding and performance tuning.
  • Own the full storage hardware lifecycle – qualification, burn‑in, firmware, OSD add/drain/replace, node retirement and refresh – and automate these processes.
  • Automate provisioning of storage nodes on OpenStack‑managed bare metal using Ironic, RAID, SEDs and VLANs, and manage artefacts such as images and certificates.
  • Run upgrades, rebalancing and reconfigurations, and evolve Infrastructure‑as‑Code and GitOps pipelines with Ansible, Terraform and FluxCD/ArgoCD.
  • Plan capacity, monitor consumption, and build testing, failure‑injection and observability around performance, durability and security.
  • Leverage AI‑assisted engineering tools (LLMs, agentic tooling) for development, testing, incident triage and automation.

Required profile

  • Several years of hands‑on experience as an SRE, Storage Engineer or Platform Engineer running production Ceph clusters.
  • Deep knowledge of Ceph fundamentals (CRUSH, pools, placement groups, replication, erasure coding, BlueStore, scrubbing, recovery).
  • Experience managing storage hardware end‑to‑end, including drives, firmware, controllers and OSD lifecycle.
  • Proficiency with Kubernetes as a control plane, operators, custom resources and GitOps, plus solid OpenStack troubleshooting skills.
  • Strong Linux and bare‑metal expertise, comfortable with Ansible, Terraform and automation workflows.
  • Practical use of LLMs or other AI‑assisted tools in daily engineering tasks.

Required skills

  • Ceph (CRUSH, BlueStore, erasure coding, RBD, RGW/S3)
  • OpenStack (Ironic, Cinder, Neutron, Nova)
  • Kubernetes, operators, GitOps
  • Ansible, Terraform, FluxCD, ArgoCD
  • Linux, KVM, bare‑metal provisioning
  • Go, Python
  • Prometheus, Grafana
  • Networking (VLAN, BGP)

What we offer

  • International, innovative environment with cutting‑edge technologies.
  • 32 vacation days (increasing with seniority) and flexible working hours with home‑office options.
  • Permanent contract with market‑based compensation and performance bonuses.
  • Employer‑funded pension, insurance package and 50 % public‑transport subsidy.
  • Annual €400 contribution toward sports activities and corporate benefits discounts.
  • Regular company events and free beverages.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec gridscale.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Le contrat proposé est un Unbefristet basé à Köln.
Source : ats:recruitee

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

Apply now →

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 1 day ago

Expires 1 month from now

9 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

gridscale

Köln