Engineering Manager, Serving/API
Positron Corporation United States · $225K–$350K/yr
Computer Hardware Manufacturing · 11-50 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
Lead and grow an engineering team responsible for the serving layer, including API fidelity, tokenization, tool calling, and model integration. Establish technical roadmaps and ensure high-performance, reliable production operations for LLM serving systems.
What they look for
Requirements
Requires demonstrated success managing engineering teams in LLM inference or ML systems with deep expertise in C++ and Python. Candidates must possess strong technical judgment in API design and the ability to translate ambiguous product demands into measurable engineering outcomes.
Benefits
Full description
Role Overview
Positron is seeking an Engineering Manager to lead Production Platform and Orchestration within our Upstack Engineering organization. This team owns the software and operating practices that provision, deploy, observe, upgrade, and reliably operate Positron systems in production. You will inherit a technically strong core team and help it grow into a durable organization capable of supporting a fleet that is expanding by several multiples.
This is a technical leadership role with real operational accountability. You will set direction, build the team, create clear ownership, and improve the systems and processes behind fleet orchestration, deployment lifecycle, observability, incident response, release automation, and production reliability. You will work closely with serving and API, model enablement, compiler and runtime, hardware, customer-facing, and data center partners.
The strongest candidate will combine systems depth with organizational judgment, moving comfortably between architecture, delivery, incidents, people development, and cross-functional planning. This description intentionally emphasizes outcomes and ownership over a fixed organizational chart. As the fleet and customer base grow, the function may develop dedicated groups for fleet orchestration and capacity, deployment lifecycle, reliability and observability, data center operations, customer production operations, and operational tooling.
About the role
We are seeking an Engineering Manager to lead Serving/API. This team owns the software between a customer's API request and the tokens Positron accelerators generate: what the model sees, and how its output becomes a correct, well-formed response. You will lead a strong team and grow it as Positron adds models, modalities, and serving capabilities.
The team owns:
- The OpenAI-compatible serving layer. HTTP endpoints, request validation, SSE streaming, usage accounting and API-spec fidelity.
- Tokenization and chat templates. HuggingFace tokenizers, chat-template rendering, and model-specific conversation formats such as OpenAI Harmony for GPT-OSS.
- Tool calling and structured output. Tool-schema handling, tool-call parsing and serialization, and grammar-constrained decoding (llguidance) for JSON and function calls.
- Reasoning and budgets. Reasoning-channel parsing, reasoning-effort controls, token limits and per-request budgets.
- Speculative decoding. The request-side plumbing for draft models and draft trees, acceptance metrics, and the policies that decide when speculation pays off.
- New modalities and frameworks. Vision-language model (VLM) input handling and integration paths with SGLang-style serving.
This is a technical leadership role with real product accountability. You will set direction, build the team, create clear ownership, and improve the systems behind API fidelity, tool calling, reasoning, speculative decoding, and new-model readiness. You will work closely with Production Platform and Orchestration, Compiler/Executor, hardware, and customer-facing partners.
What you will do
- Lead, coach, and grow a team of engineers spanning API serving, tokenization, structured generation, and model integration.
- Establish a clear technical and organizational roadmap for the serving layer: API surface, chat templates, tool calling, reasoning, budgets, and new-model support.
- Ensure API fidelity: OpenAI-compatible behavior, streaming, usage accounting, and error handling that customers and their tooling can rely on.
- Lead reasoning-model support and request budgets, including reasoning formats, reasoning-effort controls, token limits, and per-request accounting.
- Partner with Compiler/Executor on speculative decoding: draft-model integration, request-side plumbing, acceptance metrics, and policies for when speculation pays off.
- Lead the plan for vision-language model (VLM) support and SGLang interoperability, deciding what to adopt, what to build, and how it fits Positron's engine.
- Keep serving-layer overhead off the critical path for time-to-first-token and streaming latency.
- Keep the API layer independent of hardware topology as Positron adds platforms that run inference across multiple hosts.
- Work with Production Platform and Orchestration to define clear ownership of the shared host-level load-balancing layer, including where request policies and protections live.
- Build release gates (API conformance tests, tool-call and reasoning evals, regression suites) so new models and features ship with confidence.
- Collaborate with Production Platform and Orchestration to turn new serving capabilities into supportable production endpoints.
- Translate customer and business priorities into sequenced engineering work while protecting the team from reactive, unstructured requests.
- Hire thoughtfully, develop emerging leaders, and create ownership boundaries that remain effective as the organization scales.
What success looks like
First 6 months
- Build trust with the team and partner organizations; clarify ownership, decision rights, and the near-term hiring plan.
- Baseline API conformance, tool-call accuracy, serving-layer latency, and the largest sources of customer-visible defects.
- Establish release gates for API conformance and tool-calling accuracy that run on every release.
- Produce an agreed roadmap covering speculative decoding, VLM support, and SGLang interoperability alongside near-term model and customer needs.
6 to 12 months
- Grow the team and create durable ownership for API serving, structured generation, reasoning and budgets, and new-model integration.
- Ship speculative decoding in production for at least one model, with measured acceptance rates and throughput gains.
- Make API support for new models (chat templates, tool-call and reasoning formats) repeatable, with predictable turnaround.
- Develop engineers and technical leads who can independently own major serving domains.
12 to 18 months
- Support a substantially larger model catalog and feature surface without proportional growth in team effort.
- Navigate transition to new hardware platform with expanded capabilities and serving requirements.
- Demonstrate measurable improvement in API fidelity, tool-call accuracy, and time-to-first-token.
What we are looking for
- Demonstrated success managing and growing engineering teams responsible for LLM inference, model serving, ML systems, or a closely related domain.
- Systems depth in C++ and working proficiency in Python; comfort reviewing performance- and correctness-critical code.
- A record of turning ambiguous product demands into a coherent roadmap, explicit ownership, and measurable engineering outcomes.
- Ability to recruit, coach, and retain engineers across experience levels while maintaining a high technical bar.
- Clear written and verbal communication, sound prioritization, and the ability to make tradeoffs visible to technical and executive stakeholders.
- A hands-on leadership style: close enough to architecture and code to ask the right questions without becoming the team's bottleneck.
Especially valuable experience
- Contributing to vLLM, SGLang, llguidance, XGrammar, Hugging Face tokenizers, or similar projects.
- Strong technical judgment across LLM serving: tokenization, chat templates, sampling, streaming APIs, and the OpenAI Chat Completions and Responses APIs.
- Hands-on understanding of tool calling, function-calling formats, and structured output, and of how they fail in practice.
- Familiarity with open-source serving stacks such as vLLM, SGLang, or TensorRT-LLM, and the judgment to know when to adopt rather than build.
- Serving vision-language models: image preprocessing, vision encoders, and multimodal token handling.
- Supporting reasoning models and their output formats, such as OpenAI Harmony.
- Serving on GPU, FPGA, ASIC, or other accelerators, especially memory-bandwidth-bound inference.
- Customer-facing API products where engineering teams own compatibility, escalations, and release readiness.
Leadership profile
The strongest candidate will combine serving-systems depth with organizational judgment. They will be comfortable moving between API design, model integration, performance work, people development, and cross-functional planning. They will keep the team focused on correctness and compatibility as new models, modalities, and serving techniques arrive faster than any one team can absorb.
Role scope
This description intentionally emphasizes outcomes and ownership over a fixed organizational chart. As the model catalog and customer base grow, the function may develop dedicated groups for API and protocol compatibility, structured generation and tool calling, speculative decoding, and multimodal serving.
Why Join Us?
- You will build the production platform that turns purpose-built inference silicon into services customers can depend on, with direct ownership of how a rapidly growing fleet is deployed, operated, and scaled.
- You will shape both the technology and the organization from an early stage, defining the orchestration, reliability, and automation foundations that Positron will operate on for years to come.
Compensation & Benefits
The base salary range for this role is $225,000 – $350,000.
Please note that the figures provided represent the base salary range only and do not include other elements of our total compensation package, equity, or comprehensive benefits.
At Positron AI, we value the unique expertise each candidate brings. While the range above reflects our typical expectation for the position, we reserve the flexibility to exceed this range for candidates whose specialized skills, significant experience, or unique qualifications fall outside the standard scope of the role. Final offers are determined based on a variety of factors, including internal equity, and individual impact.
Benefits & Perks
We want you to do your best work and feel confident that you and your family are taken care of. That means comprehensive coverage, real time to rest, and support for your future.
Health and wellness
- Fully company-paid medical, dental, and vision insurance for you and your dependents
- Company-paid life and disability coverage, with voluntary options to add more
- Supplemental hospital, critical illness, and accident coverage available
Time off and flexibility
- Unlimited paid time off, we encourage everyone to truly unplug and recharge
- 13 paid company holidays
- Remote-first culture with a company-provided computer and home office setup
Compensation and future
- Competitive salary and equity
- 401(k) with company matching, eligible from day one
Visa Support
This position is open to candidates currently authorized to work in the U.S. We cannot provide new visa sponsorship for this role but are open to facilitating H-1B visa transfers for eligible candidates.
Equal Opportunity Employer. If you're excited about the role but don't meet every bullet, we'd still love to hear from you.
Similar roles
-
Engineering Manager, Knowledge Graph Platform
Reddit United States · $217K–$304K/yr
-
Senior Engineering Manager (REMOTE)
Fanatics Inc. New York, New York, United States · $163K–$265K/yr
-
Senior Engineering Manager (REMOTE)
Fanatics Betting & Gaming New York, New York, United States · $163K–$265K/yr
-
Lean & Industrial Engineering Manager
Siemens Pendergrass, Georgia, United States · $123K–$211K/yr
-
Advanced Manufacturing Engineering Manager
Siemens Pendergrass, Georgia, United States · $123K–$211K/yr
-
Senior Engineering Manager, Containers
Mirantis United States