Inference
Running a trained model to produce outputs. For deployed AI products it is where most of the compute bill, and most of the latency, comes from.
Hiring data as of September 20, 2026 · refreshed daily
- companies naming it
- 14 of 47companies naming it
- open roles naming it
- 37open roles naming it
- of those based in NYC
- 19of those based in NYC
- median posted base (n=13)
- $238kmedian posted base (n=13)
What Inference means
Inference is the act of using a model, as opposed to training it. Every API call to a language model is inference. For a product in production it dominates cost, because training happens once and inference happens on every request.
Inference engineering is its own specialty: batching requests to keep GPUs busy, caching repeated context, quantizing weights, choosing serving frameworks such as vLLM or TensorRT, and trading latency against throughput. The economics are moving fast enough that purpose-built inference hardware is now a category.
Inference in New York job posts
14 of the 47 New York AI companies we track name Inference in at least one open role today, across 37 postings. 19 of those are based in New York; the rest are remote or in the companies’ other offices.
This is engineering vocabulary: 68% of the postings that name it are Engineering roles.
How to read it
Counted only where the posting also names serving, latency, GPUs, or similar, to exclude the statistical sense of the word. A posting centered on inference is a systems role, strong on performance engineering and GPU behavior.
What these roles pay
Midpoints of 13 distinct salary ranges posted on New York roles that name Inference. Base salary only; on-target earnings are excluded, and a range repeated across one company’s postings counts once. For comparison, the median across all New York engineering postings is $215k. A term’s pay reflects the roles that name it as much as the skill itself — see the role mix below before reading a premium into it. Full bands by function and seniority are in the NYC salary calculator.
Trend
We began extracting this term from posting text on September 20, 2026. Job descriptions disappear when a role is filled, so the series cannot be backfilled; a trend line appears here once a week of data exists.
Who names it
Open roles naming Inference, by company.
- Modal11
- AlphaSense6
- Runway4
- Normal Computing3
- Hebbia2
- Patlytics2
- Pinecone2
- ElevenLabs1
- Headway1
- Hume AI1
- K Health1
- Mirage (fka Captions)1
Also: Ramp, SmarterDx.
Which roles
The same postings, by function and by seniority.
- Engineering25
- Research8
- Product2
- GTM1
- Other1
- Mid19
- Staff/Principal10
- Senior4
- Entry2
- Manager1
- Director1
Named alongside Inference
Terms that appear in the same postings far more often than chance would put them there. The number is how many of the 14 companies pair the two.
- Quantization and distillation 4
- GPUs 4
- CUDA 3
- Multimodal AI 4
- Fine-tuning 4
- LangChain 3
- MLOps 3
- Rust 4
- PyTorch 5
- Elasticsearch 3
Open roles naming Inference
A sample from today’s data, New York roles first and one per company before any repeats. Links go to the employer’s own posting. For the full market, see AI jobs in NYC.
- AI Engineer ↗Patlytics · New York
- Applied AI Engineer ↗Ramp · New York, NY (HQ)
- Applied Research Engineer, Agents ↗Hebbia · NYC
- Customer Engineer ↗Modal · New York
- Hardware Engineer, PCB ↗Normal Computing · New York City
- Member of Technical Staff, Robotics Engineer ↗Runway · New York, NY
- Research Engineer, Generative Video ↗Mirage (fka Captions) · Union Square, New York City
- Senior Software Engineer - Backend & Machine Learning ↗Hume AI · New York, New York, United States
On the AI in NYC Show
- Episode 24: Cerebras IPO, the Dead Internet, and the AI Backlash
Hosts only
- Episode 10: Nvidia vs. AMD and the AI Power Crunch
Doug O’Laughlin, SemiAnalysis
How this is counted
Every day we read the open roles on the public job boards of the New York AI companies in our coverage — the NYC AI 100 and the AI in NYC Show roster. This term is matched against each role’s description after removing the text a company repeats across its postings, so an “About us” paragraph cannot tag every role the company has open. A vendor’s own postings never count toward its own name.
The figure means named in a job posting. It does not mean used in production, and a “nice to have” counts the same as a requirement. Companies are the headline number because posting counts are dominated by whichever few employers are hiring hardest this month.
Full method, including the pay rules, is on the glossary index; the underlying series is the NYC AI Hiring Index.
Inference — common questions
What does Inference mean in an AI job posting?
Running a trained model to produce outputs. For deployed AI products it is where most of the compute bill, and most of the latency, comes from. Counted only where the posting also names serving, latency, GPUs, or similar, to exclude the statistical sense of the word. A posting centered on inference is a systems role, strong on performance engineering and GPU behavior.
How many New York AI companies are hiring for Inference?
As of September 20, 2026, 14 of the 47 New York AI companies we track name Inference in the description of at least one open role, across 37 postings (19 based in New York). The count refreshes daily from the companies’ own job boards.
What do roles that ask for Inference pay in New York?
Across 13 distinct posted salary ranges on New York roles naming Inference, the median midpoint is $238k base, with the middle half between $225k and $275k. The same figure for all New York engineering postings is $215k. These are employer-posted ranges required by New York’s pay transparency law, base salary only.
Which skills are asked for alongside Inference?
In the same postings, the terms most distinctively paired with Inference are Quantization and distillation, GPUs, CUDA, Multimodal AI, Fine-tuning. We rank pairings by how much more often they appear together than apart, so near-universal terms such as Python do not crowd out the informative ones.
Need someone who has actually shipped this?
A keyword in a posting is easy to match and hard to verify. We match companies with AI engineers and consultancies vetted on production work, and tell you when the project needs a different skill than the one you named.
Related Content
NYC AI Hiring Index
Weekly data on who is hiring, and what they post as pay.
AI Jobs in NYC
Open roles pulled daily from company job boards.
NYC AI Engineer Salary Calculator
Posted pay bands by level, employer type, and specialization.
AI Engineer Salary: NYC
What AI engineers earn in New York, honestly told.
How to Get Into AI
Every real path into the industry, mapped honestly.
NYC AI 100
New York’s AI companies, verified and tiered, with live roles.