Directed multimodal capture

It happened.
Nobody filmed it

Second Orbit produces multimodal training and evaluation data for teams building models that work in the physical world. We design the situation, capture it with an expert in the loop, structure the ground truth, and measure what the model learned

Our team has delivered multimodal collection, annotation and evaluation programs one week from brief to locked spec.

Google Meta

What the asks have in common

Not rare
Unrecorded

A volume vendor will sell you a thousand hours of someone cooking dinner competently. None of it reaches the forty seconds where it goes wrong — because nobody was running two cameras and an annotation schema when it did.

FOUR STAGES, ONE LOOP

A labeling vendor can’t design the situation.
A capture shop can’t tell you whether it worked.

These belong together because each one decides the next.

And then back to the beginning Because what the evaluation shows is how we decide what to capture next.

ADVISORY

Some work doesn’t fit a standard data program. AI strategy, coaching and enablement, use-case prioritization, custom applications, and collection scope that falls outside the usual shape. We take a small number of these each year.

Advisory →

Capture

Environments, not stock categories.

Every environment below is a real working setting with a domain expert on the line.
Filter by what your model has to survive.

To view more click here

The specification

You write it.
We do not have a house format.

Select what your model actually needs. The sheet on the right is the delivery specification that would be produced — it is the same document our capture leads work from.

Delivery specification0 selected

Illustrative. Volumes, rates and schema keys are agreed per programme.

How they are produced

Directed is not scripted

The fault is arranged in advance. The diagnosis is genuine. There is no second take — which is the whole point, and the reason this cannot be scraped.

CAPTURE CONFIGURATION

Every program is a designed situation with a checkable outcome. What gets designed into it depends entirely on what your model is failing at. An induced fault in one program, a wayfinding decision in the next, a handoff between two people in the one after that.

01Synchronized multi-view

Head-mounted, fixed wide, overhead, close-up — as many as the task needs, hardware-synced. Frame offset recorded, not estimated.

02A domain expert on the line

Not an actor. The person diagnosing is the person who does this for a living.

03The condition is designed

Whatever your model is failing at, we arrange for it to occur — an induced fault, a substitution, an interruption, an ambiguous instruction, a handoff between two people. What happens next is not arranged.

04Labels at frame accuracy

Action boundaries, object tracks and the verbal channel, aligned to the same clock.

05A manifest per clip

Consent, release, capture conditions, labeler identity, schema version.

06Contextual metadata

Location class, heading and dwell where the task is spatial. Device, lighting, camera placement and sync offset on every clip. Attached at capture, not reconstructed afterward.

How a program runs

Six steps, and you are in every one of them.

01Scope

You describe one behavior your model gets wrong. We come back with a written protocol — the objective, the environment, the participants, the devices, the coverage, and the conditions we intend to induce. You approve it before anyone is booked.

02Recruit

We find people who do this for a living, in the setting where they do it — a technician at their own bench, a cook in a working kitchen. Consent and release are agreed before anything is arranged, documented per contributor and recorded against every clip.

03Direct

Directed is not scripted. The condition is arranged in advance — an induced fault, a substitution, an interruption, a handoff between two people. The participant is briefed on the situation, never on the outcome. What happens next is genuine, and there is no second take.

04Capture

As many synchronized views as the task needs — head-mounted, fixed wide, overhead, close-up. We have run six on a single work area, all on one clock, with frame offset recorded rather than estimated. A domain expert watches the feed and corrects in real time.

05Annotate

Your annotation schema, not ours. Six annotation families, from pixel-level spatial labels to cross-modal alignment, with boundaries marked to the frame. Multi-pass review, gold-standard validation, and inter-annotator agreement measured per task family.

06Evaluate

We build the evaluation on the same ground truth we collected — answer-blind pipelines, human gold standards, rubric scoring. Or you run your own and tell us what it showed. Either way, what the model still gets wrong becomes the scope for the next round.

What you receive

Media, labels, transcript, manifest

Delivered to your schema, not ours. Every clip carries its own provenance record, and every label is traceable to the person who made it.

CAPTURE APPARATUS

01Media

Every view at original capture rate, no re-encode above your ceiling. Frame offset recorded per clip, so the views stay on one clock.

02Labels

Action segments, object tracks, and the failure interval marked — to the frame, against your annotation schema. Tracks held through occlusion and re-entry.

03Transcript

Diarized, timestamped and aligned to frame. The participant’s narration and the moment of correction are tagged, not left for you to find.

04Manifest

One record per clip: device, sync offset, capture conditions, consent and license terms, annotator identity and schema version. Attached at capture.

05Evaluation split

Held-out by scenario, not by random frame — so the split means something. The held-out set is agreed with you before capture, not carved out afterward.

06Quality Control

Multi-pass review against the rubric agreed with you in the pilot, gold-standard validation, and inter-annotator agreement. Sessions that miss it don’t ship.

Evaluation case study

We built the evaluation
before anyone had built the dataset

A frontier AI research team wanted to know whether today’s vision-language models can act as real-time instructors — watching someone work through a physical task and coaching them the way a human expert would. No off-the-shelf benchmark measures that. It needed purpose-built collection and a purpose-built evaluation.

Delivered by this team at The Blue Dot Labs. Client anonymized at their request.

EVALUATION · AS DELIVERED
Source videos162 Task domains3 Judged items9,600+ Native video vs frames~18 points Agreement, native videoabout 69% Total model spendunder $1,000

Operations

US capture where context matters.
Global operations where scale does.

An office and capture studio in Sunnyvale, California. On-shore capture when the contract requires it. Global operations when throughput does.

NETWORK AND OPERATIONS
Creators, United States200+ Creators, India100+ Creators, Philippines50+ Operations team30 Operating partnersJapan

QUALITY SYSTEMS

Multi-pass QA against a rubric agreed before collection · gold-standard validation · inter-annotator checks · human-in-the-loop review at every stage · a written rejection standard, so what disqualifies a session isn’t decided after the fact.

CREATORS

Our capture network is paid at rates we’re comfortable publishing.

Consent and provenance, in full →

Provenance

Every frame has a paper trail.

Directed capture means we know exactly who is in the frame, what they agreed to, and who labelled what. That is the part scraped data cannot give you.

Record What it contains Attached at
Consent and releaseNamed, scoped, and specific to the use. Withdrawal terms recorded alongside.Before capture
Capture conditionsLocation class, lighting, camera placement, sync offset, operator role.At capture
Direction recordWhat was arranged, what was briefed, and what was deliberately not disclosed to the operator.At capture
Annotator identityWho labelled each segment, against which schema version, and who reviewed it.At label
Chain of custodyEvery transfer of the media, signed. Nothing arrives without a history.Throughout

About

We’ve done this before

Second company, same team. Between us, two decades each in the two disciplines this work actually runs on: video and geospatial data at consumer scale, and global operations delivered to a written standard.

Amit Kela - LinkedIn Amit brings nearly 20 years of technology and operations consulting experience, leading global programs at Deloitte, MetLife, and Uber, including AI initiatives. He specializes in building and scaling complex operations across teams, geographies, and languages, with a focus on delivering consistent quality at scale. He holds an MBA from Wharton. Sanjeet Arora - LinkedIn Sanjeet spent over 16 years at Google, working across Apps, Geo, Area 120 and YouTube. He specializes in building new ventures, developing strategic partnerships, and transforming early-stage ideas into scalable businesses, with a focus on bringing emerging AI technologies into the real world.

Our first company together was The Blue Dot Labs, which we grew to multi-million-dollar revenue in about thirty months. This is the spinoff of our data and services business.

This is the Second Orbit

Start here

Send us the forty seconds you cannot find.

Describe one behavior your model gets wrong. We come back with a capture plan, an annotation schema and a sample — before any volume is discussed.

PRIMARY · SECONDARY · PROCUREMENT

WHAT A PILOT IS

Every program starts with a pilot, sized to the question.

What a pilot is →

WHAT WE DON’T DO

We don’t do text-only RLHF.

We don’t sell scraped or resold footage.

We don’t run unsupervised collection at volume and call it directed.

FOR CREATORS

Join the network

Second Orbit · Directed multimodal training data

hello@secondorbit.ai

Capabilities

Design the situation

The research question becomes a written protocol. Nothing is left to what happens to occur — and nothing is decided without you.

WHAT THE PROTOCOL SPECIFIES

The objective

The behavior your model gets wrong, written as something that can be filmed and checked.

The environment

A real working setting, not a set. Location class, lighting and constraints named in advance.

The participants

People who do this for a living. Consent and release agreed before anything is arranged.

The devices

Head-mounted, fixed wide, overhead, close-up — chosen as a variable rather than a default.

The coverage

How many views the task needs, and what each one is there to see.

The induced conditions

The fault, substitution, interruption, ambiguous instruction or handoff we arrange. What happens next is not arranged.

The protocol is the deliverable of this stage. You approve it before anyone is booked, and it comes back revised after the pilot.

Capabilities

Capture the evidence

Expert-led sessions, reactive and proactive. Real devices, chosen as a variable rather than a default. We have run six synchronized views on a single work area.

THE CAPTURE CONFIGURATION

Synchronized multi-view

Head-mounted, fixed wide, overhead, close-up — as many as the task needs, hardware-synced. Frame offset recorded, not estimated.

A domain expert on the line

The person who does this for a living watches the first-person feed and corrects in real time. Not an actor.

The condition is designed

Whatever your model is failing at, we arrange for it to occur — an induced fault, a substitution, an interruption, an ambiguous instruction, a handoff between two people. What happens next is not arranged.

Labels at frame accuracy

Action boundaries labeled to the frame rather than the second, against your taxonomy.

A manifest per clip

One record per clip: consent, release, capture conditions, annotator identity and schema version.

Contextual metadata

Location class, heading and dwell where the task is spatial. Device, lighting, camera placement and sync offset on every clip. Attached at capture, not reconstructed afterward.

Capabilities

Structure the ground truth

Six annotation families, from pixel-level spatial labels to cross-modal alignment. Your schema, exported — not ours.

THE SIX FAMILIES

01Spatial labels

Masks and boxes at pixel level, on the objects that matter.

02Labels

Action boundaries to the frame, against your taxonomy.

03Transcript

Held through occlusion and re-entry.

04Manifest

Time-aligned, with the correction moment tagged.

05Evaluation split

Speech, action and object tied to the same clock.

06Quality Control

Scored against the rubric agreed in the pilot.

QUALITY CONTROL

Multi-pass review against the rubric agreed in the pilot, gold-standard validation, and inter-annotator agreement measured per task family. Sessions that don’t clear the rubric don’t ship — and the rubric is written down, in your language, before capture starts.

Capabilities

Measure what the model learned

Answer-blind evaluation pipelines, human gold standards, rubric scoring across accuracy and response quality. Standalone, or alongside a collection program.

HOW THE EVALUATION IS BUILT

Answer-blind pipelines

The judge never sees the reference answer. What comes back is a score you can defend to a reviewer.

Human gold standards

Built by the same domain experts who ran the capture, on the same ground truth.

Rubric scoring

Accuracy and response quality scored separately, against the rubric agreed in the pilot.

Held-out by scenario

Not by random frame — so the split means something.

Capabilities

Advisory

Second Orbit primarily produces multimodal training and evaluation data for physical-world AI. Alongside that, we take on a small amount of advisory work each year — much of it with organizations we already know, and much of it starting as a question about what to collect.

AI strategy and roadmap

deciding which use cases are worth building, and in what order.

Executive coaching and enablement

working with leadership teams on how AI changes what their organization does and how it operates.

Use-case prioritization

assessing candidate applications against feasibility, data availability and value.

Custom application development

building the thing, where building it is the right answer.

Workflow and operations automation

putting AI into a process that already exists rather than around it.

Data-collection scope outside a standard program

collection or annotation work that doesn’t fit the usual shape.

We advise a small number of organizations each year, so we’re selective about fit. If you think there’s one, write to us.

Capabilities

Every program starts with a pilot, sized to the question.

The research question becomes a written protocol. Nothing is left to what happens to occur — and nothing is decided without you.

WHAT COMES BACK

Sized to the question

We’ve run 10–15 sessions in a week, and three rounds of three with fast feedback between rounds.

The footage

One environment, one condition, all views, fully labeled.

The protocol as revised

What we changed after the first round, and why.

The devices

Head-mounted, fixed wide, overhead, close-up — chosen as a variable rather than a default.

The annotation schema as agreed

In your language, written down before capture starts.

Our recommendation

Whether to scale — and if not, what would have to change first.

Join-the-network

Join us

We work with creators across the US, India and the Philippines on real-world capture. If that’s you, we’d like to hear from you.

200+United States
100+India
50+Philippines

WHAT THE WORK IS

Real settings, not a studio

Kitchens, workshops, warehouses, streets — wherever the task actually happens.

Your own expertise

We look for people who do the task for a living, not performers.

Paid, at published rates

Our capture network is paid at rates we’re comfortable publishing.

GET IN TOUCH

Or write to hello@secondorbit.ai We read everything, and we reply to everyone we can work with.

Error 404

Page Not Found

Catalogue

Browse the sample you need

Each video is catered to the client's exact specification.
Send us yours.

35 data types

Hover an ask to view it · click to keep it · or tab to the list
The asks 0 kept
Pick a scenario