Multimodal AI

Models that take in or produce more than one kind of data — text, images, audio, video — within a single system.

Hiring data as of September 20, 2026 · refreshed daily

companies naming it
7 of 47companies naming it
open roles naming it
13open roles naming it
of those based in NYC
9of those based in NYC
pay: under 8 posted ranges
pay: under 8 posted ranges

What Multimodal AI means

A multimodal model works across data types: reading a chart and answering in text, generating video from a description, holding a spoken conversation. In New York the term is tied to the city’s generative-media and voice companies.

Multimodal AI in New York job posts

7 of the 47 New York AI companies we track name Multimodal AI in at least one open role today, across 13 postings. 9 of those are based in New York; the rest are remote or in the companies’ other offices.

Research roles make up 54% of the postings that name it, with the rest spread across other functions.

How to read it

Concentrated in research and machine-learning engineering roles at companies whose product is generated media. It usually implies training or adapting models, not only calling them, so expect the posting to name distributed training, fine-tuning, and inference as well.

What these roles pay

Not enough data to publish. We need at least eight distinct posted salary ranges on New York roles naming Multimodal AI before showing a median, and today there are fewer. For posted pay by function and seniority across the whole market, use the NYC salary calculator or the NYC AI engineer salary guide.

Trend

We began extracting this term from posting text on September 20, 2026. Job descriptions disappear when a role is filled, so the series cannot be backfilled; a trend line appears here once a week of data exists.

Who names it

Open roles naming Multimodal AI, by company.

  • Mirage (fka Captions)3
  • Modal3
  • Runway3
  • ElevenLabs1
  • EvolutionIQ1
  • Normal Computing1
  • SmarterDx1

Which roles

The same postings, by function and by seniority.

  • Research7
  • Engineering5
  • Product1
  • Mid6
  • Staff/Principal4
  • Senior2
  • Entry1

Named alongside Multimodal AI

Terms that appear in the same postings far more often than chance would put them there. The number is how many of the 7 companies pair the two.

Open roles naming Multimodal AI

A sample from today’s data, New York roles first and one per company before any repeats. Links go to the employer’s own posting. For the full market, see AI jobs in NYC.

How this is counted

Every day we read the open roles on the public job boards of the New York AI companies in our coverage — the NYC AI 100 and the AI in NYC Show roster. This term is matched against each role’s description after removing the text a company repeats across its postings, so an “About us” paragraph cannot tag every role the company has open. A vendor’s own postings never count toward its own name.

The figure means named in a job posting. It does not mean used in production, and a “nice to have” counts the same as a requirement. Companies are the headline number because posting counts are dominated by whichever few employers are hiring hardest this month.

Full method, including the pay rules, is on the glossary index; the underlying series is the NYC AI Hiring Index.

Multimodal AI — common questions

What does Multimodal AI mean in an AI job posting?

Models that take in or produce more than one kind of data — text, images, audio, video — within a single system. Concentrated in research and machine-learning engineering roles at companies whose product is generated media. It usually implies training or adapting models, not only calling them, so expect the posting to name distributed training, fine-tuning, and inference as well.

How many New York AI companies are hiring for Multimodal AI?

As of September 20, 2026, 7 of the 47 New York AI companies we track name Multimodal AI in the description of at least one open role, across 13 postings (9 based in New York). The count refreshes daily from the companies’ own job boards.

What do roles that ask for Multimodal AI pay in New York?

We do not publish a pay figure for Multimodal AI yet. Fewer than eight New York postings naming it carry a distinct posted salary range, and below that number a median says more about two or three employers than about the market.

Which skills are asked for alongside Multimodal AI?

In the same postings, the terms most distinctively paired with Multimodal AI are Distributed training, Quantization and distillation, Reinforcement learning, Tool use, Fine-tuning. We rank pairings by how much more often they appear together than apart, so near-universal terms such as Python do not crowd out the informative ones.

Need someone who has actually shipped this?

A keyword in a posting is easy to match and hard to verify. We match companies with AI engineers and consultancies vetted on production work, and tell you when the project needs a different skill than the one you named.