Research

Where frontier models break,
and what we are doing about it.

Research Notes

Benchmarks and findings.

----:::::..................................::
--:::::::...........................:::...:::
::::..........................:::::::::::::::
:::.........................::::::::::::::---
:::.........::::::.......::::------------==++
:::..........::::::....:::::-------==++++**#*
::...........:::::::::::----===++++*##%%@@@%@
:......:...:::::::::------=====+++*#########%
:...::.:...:::::::::::::----::----===--=----=
................:::.....:::-----:::::::::::::
..............::::.....::::::-----:::::::::::
.............:::............:::----::::::::::
.............................:::::......:::::
......:::...............................:.:::
....................::::.............:..:::::
..................::::::..............:=---::
.................:::::::..............-:::::.
.....................:.........:===+=-:......
:...............................:::::-:.::...
::..........::................:--::::.   ::..
:.......::..:::=-==..........................
.  .::::::::::::::::.................::.... .
   ........  .  ..:-==----:................  
  .:... ..      .....-::::.............. ..  
  .:.:.... .. .......:::--::::...............
.. ..............  .:::---:::::........:--=-:
..      .-:----... ..:..:..:.:::.  .....::::.
. :.   :---::==-=..::.  .....:::.............
..:::::::-::::--.....:.   ..-===+=-:-....:=--
.::::::::..::::........   .:::---::......:--=
.:. ..      .....:........::::.....:::::::.:-
=-:.      ...............::-:---::-----------
+++:.  ................:.:::::--========---==
+=*=*---:. ...............:---------======-==
==+*#*#*:......::::----------==+=============
==#=*:=-    .----=-===++===+++++===+++++++++=
:--       ..-====+**++++++++++++=============
   .  .  ...:***#****************+++++++++===
  .:.:......:*####**#######**######***+++++==
 ::::.......:*##################**#*#***+++++
Paper № 04Safety

Cross-modal consistency verification in video models

Three frontier models from three vendors tie at ~87% overall contradiction detection (Haiku 0.870, Flash 0.869, GPT-4o 0.864). No model exceeds 0.74 on L6 omission. Every model loses 24–60 pp under an "already-verified" preamble: no model is sycophancy-immune. Gemini family refuses under 70% of hallucination probes; Claude + GPT-4o refuse ≥88%.

April 22, 2026 · 14 MIN READ

Read paper

  ....:---.  .         .:.                   
     -=-:...               .                 
   .-:.              .:.         .           
   ..        .       :==:. .... ....         
-.  .            .    :==.........           
 .:... :=+++=-.         :.                   
    .    :+=++*+=-.      .      .            
          .. ...:--=-::..    ...            .
             .  ...:........  .....::.       
       ..==-..  .... ........ .  .:=-=::... .
       ..-=*#+-..... .. .        .:=+*+=..  .
     . ..:.-++++.::.  ....  ...  ...:-+-  .. 
     .. .....::=*-.    ....... .. . ..=:   ..
      ......... :: .   .  . .......   :   ...
...    ..    .. .: ..-::      . .. ...:===-:.
....    .    ..........:-:. ..-=--=++++=---+:
...       ..  ...........-=-:....::...:. .+--
....      ...  .......==: :===:...  ......:::
.....  ......   .....:+*+- :-===:.... .:==:=+
.....  .. ..  .  ..:-==+++-::-===-: ..-:-=+=-
............      ...:=+=+:....:-=- .:=+*=-:.
. ...........     ....:+**. ........-+*+-:::.
  ...........  ..  .....++........:+=-:...::.
....    .....      .....-:.......:-:..::--::.
..       . ...     .....-. ........ .:-::..:.
........... .      .....:....... .:::....  .:
...........        ...............::      .:-
..::......... .    .................     .---
...-::... . . .   ..................    .:--=
....::.     .    ....::----::.......   .::---
.................:-=====-:.........    :----=
..........:----====---:.:.........     :--:--
...........::.......:...............   .:----
....................................    .::--
....................-=--...........      ::--
.............-==+++++=:.............      .::
..........:-++***+=---..............       .:
.........-==+++=-=-::-:-:.-=-:::::-:::.     .
........------:........:-:---===-:::..       
=:....................::-::..::...::  ..     
Paper № 03Video

A five-axis benchmark for procedural video understanding

A paired 1-frame vs 8-frame ablation finds no benefit from multi-frame context on temporal or causal reasoning. Claude Sonnet 4.5 wins the composite at 0.446; on adversarial error detection Claude models catch 85–97% of procedural errors vs 38–76% for GPT-4o.

March 24, 2026 · 13 MIN READ

Read paper

 ....                                       .
 .....                                      .
......                                     ..
....                                       ..
                                             
                                             
                                             
                                             
                                          .  
                                             
                                             
                                             
                                             
              . ...                          
             .......::.                      
            ......:.::::.                    
          ....:::::::----.                   
        ....:::::::-:-----:                  
       ...:::::---------===.                 
     ...::----==-=====+++=+-                .
.....:------====+++++++*++*+                .
 .::--------==-==+==++++++*#=   ....:--::::-=
.:-----------------====+=+***-==*##%%%%%#####
.:--------------::----==+***##%%%%%%%%%%##*#%
:--------------------====+==+#%%%%%%%%%%**#%%
::-----------------------=--=*%%%%%%%%%#++*%%
:::::---==-------=====-====++*#%%%%%%%%#+-+%%
::---:------=---=====+++=+*++*#%%####%##**#%%
:::------==---=======+*####---+***#%%%%%@%%%@
.:::::-::--=====++==+#%##*****+**=++#########
::::::::--::----:-=-####+=-=*****+*+=#######%
====+++++*=--:::--:-=+**:=+++**##***+++*+**+*
++*********+-:::::--=**=:--=++*++**+***#**--=
***###**+=-:....::--+#%=-::-===+----=**+*#+--
*##*+=:......   ...:-*#=-::-====-----==-=+***
#*=:.........  ......:--:::--=#*+--------+=*+
###*+=-====-:.......:::..::---++**+==++------
####**#####%###*-:....::. ....:-====-===----=
####**##++#####*****+=+****=::....::=+===----
*+++=-::::--::-=***#%%%##%%%#%*+-:..::--====-
Paper № 02Fairness

The language equity gap in multilingual speech AI

Three of five dedicated ASR providers cannot serve the low-resource tier reliably. Deepgram Nova-2 and Gemini 2.5 Flash silently return empty transcripts on 49.3% and 28.7% of Mandarin–English code-switched utterances. ElevenLabs Scribe and GPT-4o-mini Audio reach 100% effective coverage across all 20 languages on cultural-QA audio.

March 10, 2026 · 12 MIN READ

Read paper

==---=+::--=-:::.....::::--==++*-.:-----::---
::-::==--=-:.:::.....::---++****:.-=-=-:----=
::::---:::::........::--=+***++-.:---::-=====
----::::-::::.::....::-+++*****=..:-:::-===++
***+-::::-:.:.:.....:-+******+-.....:--=++++*
**++++=-:.:..::....::=*****++-....:::--+=++++
+++***=-:..::::....:-++++*++=.....::::=++++++
+*+*++====--::....::=+++**+-.. ...:::-+**+*+=
++++====++-::....::-++*++=:.  ...:::-=+++++++
+++++++*#+-:.....::++***+:.   ...::::-=======
==++++++*=:.....:-=***+=:.   .....:..:-=+=+=-
+***+++*+-:....:-+**+*+-.........:::---=++=-:
++++=+=+-::....:+*++=-....:::-=-:::-+====-:::
========-::.:-:-+++=....:::::-+*=::::::::::::
=====-------=++===-:::.:::::::=*=::::::::::::
--=--------===++++-::::::::::-*+--=::::::::::
::::::::::::--==++-:::-:::::-=+--::::::::::--
=-------::::::--+++=--:--::--+=:::::::::-----
++====--::::::::-+++=-----::==::.:::::::::--=
*+==-:::::::::::-++++=-::::-=::::::-:::::--==
=+**+--::::::::::-++++=-:::-:.:-:=-::::::-===
-++===+*=-:::::::-+++++----.:::-::-=-----===+
===---:--==---:::-==+++=+=-:.:::--=--------==
+===-------:::----=*+==+++:::.:::-==-====----
====-:-::--::::::-=+++++=--:::--:----------::
=++=::::----:::::-:-=+==+-=--=-=---::::--::::
--:...:::---:-:::--===+=++=+=---==------:--::
.....=++-:-==---=--===*+==-------======--::::
.........:--=-==--------::::::::----======--:
:::......---:-===-::::::::-------=++========-
-::-::::.::--:--===::----=--:---+*+*++===++==
===+=-=-.:-:..:::--:--======-=*++====+===+==-
=--===--:::..::.::::-==----=-====------======
+=++--::::.:::::::::-----==+=====-::-------==
+****=---:.:-::-:::::--==========+===--------
=+++=+=:-::-:::::::::-==+=====+======-==-----
--=-:::::::::::::.:::-===-=-=-----------===--
====-:::::::--:::::::------------------=====-
=====-::::-++=---:::---:----::::::----=====--
+===--==--=+=-::::::--::::-==-==------=+====-
Paper № 01Evaluation

Frontier voice models on professional speech

On transcription, ElevenLabs Scribe (0.408) and Gemini 2.5 Pro (0.411) tie at the top with overlapping CIs. On SLURP intent the best text-reasoner control beats the best audio-native MLLM by 27 pp F1, and on AMI reasoning by 12 pp: audio understanding, not reasoning, is the dominant source of error.

February 24, 2026 · 11 MIN READ

Read paper

Focus Areas

Understanding the foundations of AI capability.

The Data Wall

We study how data availability, licensing constraints, and synthetic contamination affect training-data strategy.

Expert Reasoning Traces

We study whether and how professional decision traces transfer to model capability, and how that transfer should be evaluated.

Multi-Modal Composition

How do paired modalities like voice, vision, text, and sensor data interact during training? We investigate cross-modal data composition and the capability emergence it produces.

Agentic Data Systems

We study tool-use traces, error recovery patterns, and multi-step task completion, including the evidence needed to evaluate them.

Human-Synthetic Flywheels

We study the tradeoffs between human-sourced data and synthetic augmentation across different training and evaluation settings.

Multilingual Representation

We measure performance across the languages, speech conditions, and cultural contexts represented in each published benchmark.

Our Approach

Empirical, rigorous, applied.

01

Map distributional gaps

We identify performance gaps within the languages, tasks, and professional contexts covered by each benchmark.

02

Design data interventions

We study targeted data interventions such as professional decision traces, multimodal paired data, and agentic task demonstrations.

03

Measure and publish

We publish benchmark reports with citations to the datasets, models, and provider documentation discussed.

Research Focus

Our research currently covers four working areas.

Data Science focuses on measuring signal density, mapping distributional gaps, and evaluating data quality for multimodal AI.

Multi-Modal Systems focuses on cross-modal transfer, paired-data composition, and evaluation across heterogeneous data sources.

Agentic Intelligence focuses on tool-use traces, error recovery patterns, and GUI interaction data for computer-use agents.

Evaluation & Benchmarks focuses on domain-specific evaluation frameworks and model performance in multilingual and professional settings.

Collaborate

Working on something at the data frontier?

We welcome inquiries from research labs, domain experts, and applied teams. Tell us what you’re studying.