Evals (LLM evaluation)

Repeatable tests that score a model or agent’s outputs against examples with known good answers, so a change can be judged by measurement instead of by feel.

Hiring data as of September 20, 2026 · refreshed daily

companies naming it
25 of 47companies naming it
open roles naming it
49open roles naming it
of those based in NYC
30of those based in NYC
median posted base (n=21)
$230kmedian posted base (n=21)

What Evals means

An eval is a test suite for behavior that is not deterministic. You assemble representative inputs, define what a good output looks like — an exact answer, a rubric, a comparison judged by another model or by a person — and score the system against it every time the prompt, the model, or the retrieval changes.

Evals are what let a team upgrade a model or rewrite a prompt without guessing whether things got better. Teams that lack them tend to ship on the strength of a few hand-checked examples and find the regressions in production. Building the dataset is usually the expensive part, because it requires deciding, case by case, what correct means for the product.

Evals in New York job posts

25 of the 47 New York AI companies we track name Evals in at least one open role today, across 49 postings. 30 of those are based in New York; the rest are remote or in the companies’ other offices.

Engineering roles make up 59% of the postings that name it, with the rest spread across other functions.

How to read it

One of the clearest signals of a team that runs AI in production rather than in demos. It appears in engineering, research, and product postings, and its most distinctive companions in the same postings are tool use and RAG. If a posting lists evals, bring an example of a metric you defined and a decision it changed.

What these roles pay

$180k25th percentile
$230kMedian
$260k75th percentile

Midpoints of 21 distinct salary ranges posted on New York roles that name Evals. Base salary only; on-target earnings are excluded, and a range repeated across one company’s postings counts once. For comparison, the median across all New York engineering postings is $215k. A term’s pay reflects the roles that name it as much as the skill itself — see the role mix below before reading a premium into it. Full bands by function and seniority are in the NYC salary calculator.

Trend

We began extracting this term from posting text on September 20, 2026. Job descriptions disappear when a role is filled, so the series cannot be backfilled; a trend line appears here once a week of data exists.

Who names it

Open roles naming Evals, by company.

  • Normal Computing6
  • Headway5
  • AlphaSense3
  • Rogo3
  • Anterior2
  • ASAPP2
  • Clay2
  • Comet2
  • ElevenLabs2
  • K Health2
  • Mirage (fka Captions)2
  • Patlytics2

Also: Pinecone, Ramp, SmarterDx, Alloy, Arthur, EliseAI, Hebbia, Maven Clinic, Rillet, Runway, Sixfold, Slingshot AI, Tennr.

Which roles

The same postings, by function and by seniority.

  • Engineering29
  • Research8
  • Product5
  • Operations2
  • GTM2
  • Other2
  • Mid26
  • Staff/Principal11
  • Senior9
  • Entry3

Named alongside Evals

Terms that appear in the same postings far more often than chance would put them there. The number is how many of the 25 companies pair the two.

Open roles naming Evals

A sample from today’s data, New York roles first and one per company before any repeats. Links go to the employer’s own posting. For the full market, see AI jobs in NYC.

How this is counted

Every day we read the open roles on the public job boards of the New York AI companies in our coverage — the NYC AI 100 and the AI in NYC Show roster. This term is matched against each role’s description after removing the text a company repeats across its postings, so an “About us” paragraph cannot tag every role the company has open. A vendor’s own postings never count toward its own name.

The figure means named in a job posting. It does not mean used in production, and a “nice to have” counts the same as a requirement. Companies are the headline number because posting counts are dominated by whichever few employers are hiring hardest this month.

Full method, including the pay rules, is on the glossary index; the underlying series is the NYC AI Hiring Index.

Evals — common questions

What does Evals mean in an AI job posting?

Repeatable tests that score a model or agent’s outputs against examples with known good answers, so a change can be judged by measurement instead of by feel. One of the clearest signals of a team that runs AI in production rather than in demos. It appears in engineering, research, and product postings, and its most distinctive companions in the same postings are tool use and RAG. If a posting lists evals, bring an example of a metric you defined and a decision it changed.

How many New York AI companies are hiring for Evals?

As of September 20, 2026, 25 of the 47 New York AI companies we track name Evals in the description of at least one open role, across 49 postings (30 based in New York). The count refreshes daily from the companies’ own job boards.

What do roles that ask for Evals pay in New York?

Across 21 distinct posted salary ranges on New York roles naming Evals, the median midpoint is $230k base, with the middle half between $180k and $260k. The same figure for all New York engineering postings is $215k. These are employer-posted ranges required by New York’s pay transparency law, base salary only.

Which skills are asked for alongside Evals?

In the same postings, the terms most distinctively paired with Evals are Tool use, RAG, Prompt engineering, Fine-tuning, Inference. We rank pairings by how much more often they appear together than apart, so near-universal terms such as Python do not crowd out the informative ones.

Need someone who has actually shipped this?

A keyword in a posting is easy to match and hard to verify. We match companies with AI engineers and consultancies vetted on production work, and tell you when the project needs a different skill than the one you named.