The methodology · Published on purpose

How we vet AI engineers: the bar, published

Every talent network claims “top 1%.” Almost none will tell you what the phrase means or how it is measured — because for most of them it is a marketing ratio, not a bar. Here is ours, in full: eight disciplines, four levels, scenario-based assessment with a live defense, scored by practitioners who have shipped the thing they are scoring.

Why we publish the bar

A credential you cannot inspect is just an adjective. Clients deserve to know what “vetted” bought them, and engineers deserve to know what they are being measured against before they invest hours in it. So we publish the modules, the rubric levels, and the process.

What stays private: the scenarios themselves and the scoring anchors, which rotate. Not because secrecy makes the bar higher, but because leaked scenarios turn a judgment exercise into a memorization exercise. The design principle throughout is simple: judgment can’t be crammed. A bar that collapses when candidates prepare for it was measuring preparation, not ability.

The eight disciplines we assess

No candidate is assessed on all eight — packs are assembled per engagement from the disciplines the actual work demands, typically two or three. What each module scores:

M1RAG & retrieval systems

Failure-mode fluency; chunking, embedding, and reranking tradeoffs; the instinct to measure before fixing.

M2Evals & quality

Designing evaluation suites under real budgets; golden-set construction; knowing that evals are not benchmarks.

M3Fine-tuning & adaptation

Judgment between prompting, retrieval, and fine-tuning; realism about data requirements and cost.

M4Agentic systems

Decomposition judgment; containing failures; knowing when to remove autonomy rather than add it.

M5LLM security & safety

Prompt-injection and data-leakage literacy; mitigations that are practical rather than theatrical.

M6AI infrastructure & cost

Caching, routing, and model-tiering fluency; cutting spend without silently cutting quality.

M7Data engineering for AI

Data-quality instincts; pipeline pragmatism when the source data is messy — which it always is.

M8Product judgment & scoping

When not to use an LLM; shipping the six-week version; the maturity to refuse to build the wrong thing.

The four levels — and what “top 1%” actually means here

L1 · Aware

Correct vocabulary, no scars. Can discuss the technology; hasn’t been responsible for it.

L2 · Competent

Sound decisions with guidance. Has shipped adjacent work; most working engineers land here — it is a respectable level, not a failure.

L3 · Production-proven

Anticipates failure modes before they happen, because they have owned one in production. Evidence required, not asserted.

L4 · Expert

Reframes the problem. The person other senior engineers call when it breaks — and the only level allowed to assess others at L3.

The network admission bar: production-proven (L3 or higher) in at least two modules, and at least competent (L2) in every module claimed. When we say top 1%, this is the operational definition behind it: L3-level judgment, verified by an L4 assessor — not a percentile of an applicant funnel, which any network can inflate by widening the funnel.

The distinction matters because the levels describe scars, not knowledge. The difference between L2 and L3 is not another course — it is having owned a production failure and carrying the anticipation that ownership builds. That is what clients are paying to borrow.

The process: scenarios, then a live defense

Step one: the async pack. Two scenarios from the relevant modules, roughly 45–60 minutes each, done on your own time with whatever tools you normally use. These are judgment exercises, not puzzles — the style, illustratively: “This workload costs $40k/month. Cut it 60% without measurable quality loss — show the plan and what you’d measure to prove nothing broke.” There is no single right answer; there are better and worse ways to think.

Step two: defend your design. Thirty minutes, live, with an assessor who has shipped in the module being tested. They probe the work you submitted: why this tradeoff, what breaks first at 10× load, what you’d cut if the budget halved. AI-assisted preparation is expected and allowed — the live session exists precisely because of it. Tools draft; people defend.

Scoring. Human, rubric-based, with written evidence for every level assigned. No auto-grading, no keyword matching — the things that made traditional screening tests easy to run are the things that made them stop measuring anything.

What clients actually receive

Every match arrives with a one-page Skill Report: a plain-English verdict, the level assessed in each module against what the brief requires, two or three direct evidence highlights from the assessment, and — always — a gaps section stated plainly. A report with no gaps is sales copy, and clients learn to distrust it. Ours closes with a recommended focus for the client’s own final interview: we don’t just filter candidates, we brief the interviewer on how to spend their one conversation.

Consultancies get the team version: references we called ourselves, a review of a real reference architecture, the delivery lead assessed directly, and a team-depth map. If you are on the hiring side, the companion guide is how to hire AI engineers — the same bar, seen from your side of the table.

What this replaces

The screening tools most of the industry still runs on stopped working, each for its own reason. Resume keyword screens measure vocabulary, and vocabulary is free now — every resume says RAG, agents, and evals. Coding platforms measure the production of known answers under time pressure, which is the one task LLMs are categorically better at than people. Take-home projects measure free time as much as skill, and reviewing them well costs more effort than most teams spend. Pedigree — the brand-name employer on the resume — measures a hiring decision someone else made years ago, about a different job.

Scenario judgment plus a live defense is more expensive per candidate than any of those. That is the point: the cost sits with us, once, instead of with every client’s interview loop, repeatedly — and what it measures is the thing the client is actually buying.

How to prepare — and why cramming won’t help

You cannot study for a judgment assessment the way you study for a certification, and that is deliberate. What you can do is collect your evidence: the systems you have shipped, the metrics they moved, what broke in production and what you did about it, what it cost to run and how you found out. Engineers who own that history walk through the assessment; engineers reconstructing it from memory struggle at the live defense.

If you are earlier in the journey, the path to L3 runs through production reps, not preparation: the roadmap lays out the ladder, the interview questions guide shows what the bar sounds like in conversation, and every path into AI covers the routes in if you are not there yet.

Think you clear it?

If you read the bar above and thought “that’s describing my last two years,” you are who the network exists for. Get assessed once; carry the credential into every match.

Frequently asked questions

Can I use AI tools during the assessment?+

Yes — we expect you to. The async scenarios mirror real work, and real work is AI-assisted now. That is exactly why the assessment ends with a live session probing the work you submitted: tools can draft an answer, but they cannot defend a design under questioning. If the judgment is yours, the conversation shows it in minutes.

What happens if I assess at L2?+

Nothing bad — L2 means you make sound decisions with guidance and have shipped adjacent work, which describes most good working engineers. It is below our network admission bar, but the assessment tells you precisely which modules to strengthen and what evidence would demonstrate it. Engineers reassess after building that evidence; judgment is trainable through production experience, and the gap between L2 and L3 is usually one owned failure away.

Do consultancies go through this too?+

Yes, with a variant built for teams: we call their references ourselves rather than reading a case-study PDF, review a real reference architecture, assess the delivery lead directly on evals and scoping at minimum, and map team depth so a client knows who actually shows up after the sales call.

How is this different from a coding test?+

Coding platforms measure whether you can produce a known answer under time pressure — which LLMs do better than any human, making those tests obsolete as a signal. We assess judgment on open-ended scenarios with no single right answer, scored by practitioners against a rubric, then pressure-tested live. It is closer to an architecture review than an exam.

Who sees my results?+

Nothing goes to a client without your agreement — introductions are double-opt-in, and the report travels only when you choose to pursue a specific match. Your results are a credential you control, not a database entry companies browse.