Schonfeld logo
Posted 1h agoNew York, New York, United States

Site Reliability Engineer

MiddleOn-site (New York)Salary undisclosed
Required Skills
PythonAWSKubernetesPostgreSQLMySQLCI/CDLLMs
Job Description

The Role

We’re looking for a Site Reliability Engineer to join our AI Technology team and play a critical role in ensuring the stability and scalability of the firm’s internal Agentic AI platform. This is a high-impact, hands-on engineering position where you’ll contribute code, automation workflows, observability, metrics, and support users as issues arise. You’ll work in a dynamic, fast-paced environment where product requirements evolve, and business stakeholders are deeply engaged.

What you’ll do

  • You will set the reliability standards for Enterprise AI, defining Service Level Objectives (SLOs), error budgets, and custom incident response runbooks.
  • You will own the observability, incident response, reliability, and scalability of our AI platform.
  • Ensure that our agents, gateways, LLM proxies, and RAG pipelines operate with high availability, accuracy, and financial efficiency.
  • Support users in a dedicated help channel when issues arise — investigating root causes, providing solutions, monitoring the status of upstream dependencies, and communicating updates back to users.
  • You’ll also be involved in the team’s code quality and best practices, identify gaps in development lifecycles, and continuously improve both the core

Similar Openings in DevOps & Cloud

View all in category