Founding Inference Infrastructure Engineer
Compensation :Indicative of San Francisco market benchmark: approximately $250,000 annual base salary.
Location: San Francisco
Sector: AI

This benchmark is provided for context and does not represent the employer’s confirmed salary range. Compensation and any equity package will be confirmed during the recruitment process.

Build and own the infrastructure behind production LLM inference on next-generation AI hardware.

Our client is an AI infrastructure company building a cloud platform for specialised inference accelerators. It deploys and operates hardware designed for demanding AI workloads, helping customers improve inference speed, efficiency and cost.

Customers include AI research organisations, growing AI application businesses and cloud providers.

About the Role

Getting a model running correctly and efficiently on specialised silicon is only half the challenge. The other half is serving it reliably in production.

You’ll build and own the inference layer between a validated model and a live customer request, covering request scheduling, batching, KV-cache management, autoscaling and the failure modes that emerge under real traffic.

This is a founding engineering role within a small team, with broad scope and direct ownership. You’ll design the serving architecture, operate it in production and take responsibility for its reliability, including on-call support.

The opportunity is to build a serving stack designed specifically for the underlying hardware, with meaningful influence over architecture, technical direction and the team’s development.

What You’ll Do

Own the inference serving stack end-to-end. Design and build request routing, batching, scheduling and autoscaling across model-serving replicas.
Reduce cost per token. Optimise batching strategies, KV-cache management and hardware utilisation to improve throughput and efficiency.
Build reliability from day one. Establish monitoring, alerting and failover, and respond to incidents affecting customer workloads.
Partner with compiler and model bring-up engineers. Define the interface between a model that is correctly compiled and validated and one that is serving live requests efficiently.
Shape the serving roadmap. Help prioritise capabilities such as multi-tenant isolation, speculative decoding and new scheduling strategies.
Set the technical standard. Make architecture and code-quality decisions that will guide the serving team as it grows.

What We’re Looking For

5+ years’ experience building and operating production infrastructure, ideally including high-throughput or low-latency serving systems.
Direct production experience with LLM inference serving, including request batching, KV-cache management, continuous batching or related techniques.
Experience owning reliability and participating in on-call support for business-critical systems.
Strong systems fundamentals across concurrency, networking and scheduling, with the ability to reason about performance at the hardware level.
A self-directed approach and comfort working with ambiguity, making technical decisions and establishing practices within an early-stage team.

Nice to Have

Experience serving models on non-NVIDIA accelerators, such as TPU, Trainium/Inferentia or other specialised AI hardware.
Familiarity with vLLM, TGI, TensorRT-LLM or SGLang, including an understanding of their limitations.
Experience building and operating infrastructure at a small company or as a founding or early engineer.
Exposure to capacity planning or fleet management for specialised hardware.

US Work Authorisation

Applicants must be legally authorised to work in the United States without employer sponsorship. This position does not offer visa sponsorship, including H-1B transfers, now or in the future.

Applications are welcomed from all qualified candidates who meet these work-authorisation requirements, regardless of citizenship or national origin.

Please email your CV to  info@alfa-executive.com quoting ref JA/ALFA/HUE