
Keep our global cloud platform reliable, scalable and resilient.
Do you enjoy solving complex production challenges before they become incidents? Are you the kind of engineer who automates repetitive work, improves reliability through engineering, and believes that every outage is an opportunity to build a better system?
We're looking for a Site Reliability Engineer to join the Global Platform Team at HeadFirst x Impellam Group. In this role, you'll help build and operate the cloud platform behind our global Workforce-as-a-Service (WaaS) ecosystem, working alongside Cloud Engineers, Platform Engineers, Data Engineers and AI specialists to improve platform resilience, reduce operational overhead and ensure our Azure-based engineering landscape remains reliable, scalable and resilient.
Your impact
As a Site Reliability Engineer, your focus is simple: keep our platforms healthy, reliable and easy to operate. You'll help build the engineering foundations behind our Headless Data Architecture (HDA), running on Azure and Databricks, as well as the Custom Apps Infrastructure (CA) that powers integrations, internal applications and operational workflows across our international organisation.
Rather than spending your days reacting to incidents, you'll focus on preventing them through automation, observability and reliability engineering. You'll reduce operational toil, improve platform resilience and build systems that scale, recover automatically where possible, and give engineering teams across Cloud, Data and AI the visibility they need to operate production workloads with confidence.
What you will do
Improve the reliability, availability and performance of our Azure platform and production environments;
Build and improve monitoring, logging and alerting using Grafana, OpenTelemetry, Azure Monitor and Log Analytics;
Automate operational tasks and eliminate repetitive manual work using Infrastructure as Code and scripting;
Design self-healing capabilities and automated remediation to reduce incidents and improve recovery times;
Investigate production incidents, perform root cause analyses and implement long-term improvements;
Define, measure and improve Service Level Indicators (SLIs) and Service Level Objectives (SLOs);
Optimise platform performance, scalability and operational efficiency;
Work closely with Cloud Engineers to improve platform architecture, resilience and security;
Support Data and AI teams by improving the reliability of Azure Databricks environments;
Drive engineering best practices around observability, automation and operational excellence;
Continuously look for opportunities to reduce operational complexity and improve developer productivity.
About the role
As part of the Global Platform Team, you'll work alongside engineers across Cloud, Data and AI to improve the reliability of our Azure-based platform. Using technologies such as Kubernetes, Terraform, Databricks, GitHub Actions, Grafana and OpenTelemetry, you'll help ensure our global Workforce-as-a-Service ecosystem remains reliable, scalable and resilient.
About HeadFirst Group x Impellam Group
HeadFirst Group x Impellam Group is one of Europe's leading providers of workforce and talent solutions. Operating across multiple countries, we're transforming into a cloud-native, AI-powered organisation that connects people, technology and data through a modern digital platform. The Global Platform Team is at the heart of that transformation, enabling engineering teams across Cloud, Data, AI and Software Engineering to build and operate scalable solutions for the future.
Interested?
If you're excited about building reliable, scalable cloud platforms and enjoy solving complex engineering challenges, we'd love to hear from you. Apply today and let's discover how you can make an impact as part of our Global Platform Team.
You're passionate about building reliable systems and solving operational challenges through engineering rather than manual intervention. You enjoy understanding how distributed systems behave, thrive in cloud-native environments and are always looking for ways to improve automation, resilience and observability. You stay calm under pressure, take ownership of problems and enjoy collaborating with others to continuously improve the reliability of the platform.
Ideally, you also bring:
4+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps Engineer or Cloud Engineer;
Strong hands-on experience with Microsoft Azure;
Experience with Infrastructure as Code using Terraform;
Experience building and maintaining CI/CD pipelines using GitHub Actions or Azure DevOps;
Experience with observability tooling such as Grafana, OpenTelemetry, Azure Monitor or Log Analytics;
Strong scripting skills using Python, Bash or similar languages;
Experience supporting distributed cloud platforms in production;
Experience with incident management, root cause analysis and post-incident improvements;
Familiarity with GitOps principles and modern deployment practices;
Experience with Azure Databricks is a strong advantage;
Experience with SnapLogic or similar integration platforms is a plus.
We know the perfect candidate doesn't exist. If this role excites you but you don't meet every single requirement, we'd still love to hear from you. We're just as interested in your potential, mindset and ambition as we are in your experience.
Any questions?
Feel free to ask, I'm happy to help!
Submit your application
We’ll reach out within 48 hours to you
First interview
Assessment
Second interview
Job offer
Welcome to HeadFirst Group!
Én de mogelijkheid tot bij- of verkopen van dagen. Wij werken hard, maar vergeten niet te ontspannen.
Met meer dan 450 collega's hebben wij altijd iets te vieren. Als één team staan wij stil bij verjaardagen, jubilea en andere successen!
Werk aan je mentale gezondheid met toegang tot het platform OpenUp. Wij leren elke dag en daarom kun je bij ons gratis opleidingen en cursussen volgen om jezelf te blijven ontwikkelen.
Je doelen behalen en dit terugzien in de vorm van een bonus? Dat kan bij ons! Hard werken wordt beloond.
Werk op één van onze 4 locaties, thuis, of op elke andere plek waar jij je prettig voelt. Uiteraard ontvang jij een mobiliteits- en thuiswerkvergoeding vanuit ons.
Maak gebruik van de gratis sportfaciliteiten, fijne werkplekken en geniet van een heerlijk lunchbuffet. Een geduchte concurrent van thuis!
Erix Santman
Product Manager
Diederick van Tellingen
UX Designer
Mattijs Wassenburg
Managing Director
Christine Koekkoek
Product Owner
Voor deze opdracht dien je een bieding te plaatsen op Striive. Striive is het grootste opdrachtenplatform van de Benelux waar jaarlijks meer dan 20.000 opdrachten gepubliceerd worden.