Site Reliability Engineer
MiddleOn-site (New York)Salary undisclosed
Required Skills
PythonAWSKubernetesPostgreSQLMySQLCI/CDLLMs
Job Description
The Role
We’re looking for a Site Reliability Engineer to join our AI Technology team and play a critical role in ensuring the stability and scalability of the firm’s internal Agentic AI platform. This is a high-impact, hands-on engineering position where you’ll contribute code, automation workflows, observability, metrics, and support users as issues arise. You’ll work in a dynamic, fast-paced environment where product requirements evolve, and business stakeholders are deeply engaged.
What you’ll do
- You will set the reliability standards for Enterprise AI, defining Service Level Objectives (SLOs), error budgets, and custom incident response runbooks.
- You will own the observability, incident response, reliability, and scalability of our AI platform.
- Ensure that our agents, gateways, LLM proxies, and RAG pipelines operate with high availability, accuracy, and financial efficiency.
- Support users in a dedicated help channel when issues arise — investigating root causes, providing solutions, monitoring the status of upstream dependencies, and communicating updates back to users.
- You’ll also be involved in the team’s code quality and best practices, identify gaps in development lifecycles, and continuously improve both the core
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.