Software Engineer, ML Infrastructure
Senior ML Infrastructure Engineer
About Rebar
Rebar is building the AI operating system for commercial HVAC, Electrical, and Plumbing.
Over the past year our quoting platform has processed tens of thousands of projects across North America and we’re continuing that growth. Our customers include many of the top firms in the industry. Some of these companies are running billion dollar construction projects on workflows that still look like it's 1985.
Construction is 10% of GDP and still massively underserved by software. We are changing that.
We recently raised a $14M Series A from leading construction tech investors and are entering our next phase of growth. We are building a set of AI native products that will define how this industry operates.
About the Engineering Org
We're in the age of ai and the role of the engineer is rapidly changing. We're aware. We're being very intentional of ensuring we adapt with it. We are fostering an engineering culture of growth and development. We strongly emphasize care of craft and winning together. Everyone operates like an owner, we find a way, and we win together.
About the Role
We're hiring a Software Engineer to own and expand the platform our ML engineers depend on to rapidly iterate, experiment, and ship models. This work will span feature pipelines, training infra, evaluation, deployment, and monitoring. You should be well versed with GPU architecture and performance focused - understanding what is blocking the engineers as well as our MFU. You'll be joining a small group of talented engineers focused on delivering practical, production-ready ML systems in a fast-moving startup context.
The role is ideal for someone who is obsessive over the developer experience of the engineers they support, is interested in training, and meticulous over performance. Our team is still (somewhat) lean and engineers wear many hats. Your purview will span the whole ML lifecycle.
Responsibilities
Platform and Developer Experience
You'll help build services that act as the single front door to our ML platform
Infrastructure Consolidation and Integration
Smooth the seams for GPU inference and our cloud stack
Right now, we utilize AWS + Temporal + Modal all in tandem, but some of this has some friction and sticking points. You will optimize and refine our architecture
Observability and Operations
We utilize DataDog for our observability. You should be comfortable extending DD and Modal support to have clear GPU util metrics as well as clean insights into our entire inference flow
Collaboration and Roadmap
You will work closely with ML engineers to understand their workflows, turn one-off scripts into self-serve platform features, and participate in architecture and roadmap decisions.
What We're Looking For
You should feel confident designing developer-facing APIs and SDKs, integrating disparate cloud and SaaS services into coherent systems, and obsessing over the experience of the engineers who use what you build.
We're seeking someone with strong platform-engineering instincts who enjoys turning fragmented workflows into products teams actually want to use. This role is a great fit if you have taste in abstractions, opinions about developer experience, and a track record of making ML or data teams meaningfully more productive.
Qualifications
6+ years of experience building production backend systems, with significant time on internal developer platforms, ML platforms, or integration-heavy infrastructure work.
Strong Python + PyTorch, including profiling and debugging below the model code
3+ years of experience with cloud infrastructure (AWS preferred)
Hands on GPU performance and inference work. Ideally you have ptimizations in production: batching, mixed precision, quantization, ONNX Runtime/TensorRT or torch.compile
A proven track record operating inference at large scale across a range of model types - detection, segmentation, recognition, and LLM/VLM workloads.
Experience managing a model zoo / model registry - versioning, promotion, and governance of models from experiment to production.
Nice to Have
Experience integrating common ML tooling - experiment trackers (W&B, MLflow), feature stores, model serving frameworks - into broader platforms.
Experience with DAG / workflow orchestration frameworks such as Temporal, Prefect, or Apache Airflow.
Built a Backstage-style internal developer portal or comparable internal platform.
Familiarity with GPU compute providers (AWS, Lambda Labs, CoreWeave, RunPod).
Some ML practitioner background — you've trained or deployed models yourself and understand the workflow from the user's side.
Experience with deployment and monitoring pipelines for ML systems.
Compensation and Benefits
Salary: Competitive base salary
Equity: Meaningful equity package, commensurate with experience
Benefits: Comprehensive medical, dental, and vision coverage
Perks:
agentic tooling budget
lunches provided, dinners provided (after a set time)
great culture and office banter
This is a salaried, onsite role located in New York City's Flatiron district. We are still a startup! We love working onsite together and believe strongly that this gives us for creative problem-solving, and building strong connections. You'll be at the heart of our fast-paced operations, actively contributing to a culture that values engagement, growth, and teamwork.
