Site Reliability Engineer (SRE) - Network Products
MiddleOn-site (Paris)€60,000 – €80,000 / yr
Required Skills
PythonGo (Golang)RustNext.jsCI/CDLinux
Job Description
OUR STORY:
Join Scaleway and shape the sovereign cloud of tomorrow !
Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies.
Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector.
With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants.
Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen.
Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.
WHY WE NEED YOU ?
Our growth is driving us to strengthen our Network SRE Products team to ensure the high reliability, performance, and scalability of our storage platforms.
Your mission will be to automate, monitor, and improve the reliability, performance, and scalability of our infrastructure. You will maximize availability, optimize fault tolerance, and reduce operational overhead—ensuring robust and efficient systems for our products and services.
YOUR FUTURE TEAM
We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together. You will be part of a team of Site Reliability Engineers reporting to a Lead SRE and integrated into the SRE Guild, a collective focused on fostering best practices across engineering.
The team collaborates daily with Dev, Product, and Ops teams to improve resiliency, support service scalability, and ensure a seamless customer experience across our network solutions.
YOUR DAILY ROUTINE
- Develop automation tools and frameworks to streamline infrastructure management
- Build and maintain CI/CD pipelines using Infrastructure as Code best practices
- Implement and refine monitoring and alerting systems (OpenMetrics, OpenTelemetry)
- Ensure system reliability through incident response and root cause analysis
- Collaborate with developers and product teams to bake resilience into network systems
- Participate in architecture reviews and provide SRE perspective early in the design
- Apply principles of fault-tolerance, load balancing, and energy efficiency optimization
- Share knowledge within the team and broader engineering org via the SRE Guild
- Contribute to the reliability and performance
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.