Efficient Computer logo
Posted 2h ago•San Francisco, Bay Area OR Pittsburgh, PA

Senior Staff Software Engineer, Linux Runtime & Drivers

staffOn-site (Bay Area OR Pittsburgh)Salary undisclosed
Required Skills
PythonNext.jsLinuxPyTorchTensorFlow
Job Description

Efficient Computer is rethinking computing from the ground up to solve one of AI's biggest challenges: energy. Our Fabric architecture delivers a 10–100x improvement in energy efficiency for general-purpose computation, including AI, while supporting familiar programming languages and software frameworks—unlocking new possibilities for autonomous machines, intelligent infrastructure, wearables, and space systems. With our first processor, Electron E1, now shipping in volume and more than $97 million in Series B financing recently announced, we're accelerating customer deployments and scaling our architecture to data center performance. Join us to build the hardware and software that bring intelligence wherever it's needed and help shape the next generation of computing.

Efficient Computer is building the next generation of energy efficient, general-purpose processors designed around the Fabric architecture, a novel spatial dataflow design spun out from decades of fundamental SOC research at CMU.

We are looking for a Senior Staff Linux Runtime Software Engineer to own the user-mode components that interface with the Fabric and support the effcc compiler.  This includes owning the user-mode driver, interfacing with our kernel mode driver,  as well as defining the HAL interface exposed by the Fabric ABI.  You will be working within pre-silicon environments, involving QEMU, functional simulation, emulation, and FPGAs to ensure applications run on day one.   

This is a key role for the System Software team.   As such, you will work closely with the compiler, kernel driver, architecture, micro-architecture, and verification teams during all phases of development.   

Key Responsibilities

Own the Runtime Library and User-Mode Driver

  • Design, implement, and maintain the user-mode runtime driver for the Fabric accelerator.  This is a critical shared library which sits between the compiler and the Fabric ABI.  This library implements context management, program loading, resource management, command queue construction and submission, descriptor allocation and setup, streaming, completion queue/event synchronization, error handling,  logging, memory allocation, as well as general library support.
  • Design and implement all user-mode interfaces between the compiler, the Fabric ABI, and the kernel mode driver.
  • Define the in-memory representations generated by the compiler and consumed by the Fabric.
  • Design and implement the underlying runtime library support and interfaces needed by 3rd party ML framework backends used by PyTorch, ONNX, and/or TensorFlow.
  • Lead the user-mode efforts to bring up on first silicon.  This includes being comfortable with functional simulation, emulation and FPGA environments.
  • Establish runtime test coverage, including API unit and conformance tests, emulation-based regression suites, performance benchmarks, and hardware-in-the-loop validation.

 

Required Qualifications

  • B.S. in Electrical/Computer Engineering, Computer Science, or equivalent; M.S. a plus.
  • 7+ years of systems software experience, including significant hands-on Linux user-mode driver as well as language runtime library development for accelerators (GPUs, NPUs, DSPs, FPGAs, or similar).
  • Experience building or contributing to a production accelerator runtime software stack, from the driver API up through the user-facing runtime library.
  • Solid understanding of accelerator memory models, in particular unified memory systems, allocators, and zero-copy data paths.
  • Experience designing public APIs exposed by runtime libraries: ABI stability, versioning, error models, thread safety, and documentation.
  • Experience working with compiler teams on binary formats, loaders, launch ABIs, and kernel metadata, for ahead-of-time or just-in-time compilation.
  • Experience integrating runtimes with language runtimes or ML frameworks: Python bindings, framework backends or execution providers, and C/C++ application interfaces.
  • Hands-on experience with QEMU, emulation, or simulation environments during pre-silicon development.
  • Experience designing interfaces between kernel and user space: ABI and HAL design, synchronization, memory sharing, and error handling.
  • Strong user-mode debugging skills across crash analysis, sanitizers, tracing, profiling, and race and synchronization diagnosis.
  • Working knowledge of ARM/AArch64 Linux systems and the Linux driver model.
  • Expert C and C++ and strong Python skills; comfortable working with AI-assisted development tools for code, analysis, and debugging.

 

Desired Qualifications

  • Experience contributing to CUDA Runtime  (cudart), ROCm (HIP, ROCr, HSA), the Tenstorrent runtime (TT-Metalium), oneAPI Level Zero, OpenCL, or another accelerator runtime software stack.
  • Familiarity with LLVM or MLIR, and with how compilers for accelerators hand off to runtimes.
  • Knowledge of dataflow, spatial, or other novel compute architectures, and of how compilers and runtimes target them.
  • Experience on a first-generation silicon product, where the software stack and the hardware matured together.
  • Experience with FPGA prototyping and hardware emulation platforms (Palladium, Zebu, Veloce, or similar).

Why Join Efficient?

Build what's next. At Efficient, you'll tackle big challenges with room to take ownership and make an impact. We back you with competitive pay, equity, a 401(k) match, company-paid benefits, paid parental leave, and flexibility—plus opportunities to grow your skills and career as we scale.

Similar Openings in Backend

View all in category➔