Research
Research Notes
----:::::..................................:: --:::::::...........................:::...::: ::::..........................::::::::::::::: :::.........................::::::::::::::--- :::.........::::::.......::::------------==++ :::..........::::::....:::::-------==++++**#* ::...........:::::::::::----===++++*##%%@@@%@ :......:...:::::::::------=====+++*#########% :...::.:...:::::::::::::----::----===--=----= ................:::.....:::-----::::::::::::: ..............::::.....::::::-----::::::::::: .............:::............:::----:::::::::: .............................:::::......::::: ......:::...............................:.::: ....................::::.............:..::::: ..................::::::..............:=---:: .................:::::::..............-:::::. .....................:.........:===+=-:...... :...............................:::::-:.::... ::..........::................:--::::. ::.. :.......::..:::=-==.......................... . .::::::::::::::::.................::.... . ........ . ..:-==----:................ .:... .. .....-::::.............. .. .:.:.... .. .......:::--::::............... .. .............. .:::---:::::........:--=-: .. .-:----... ..:..:..:.:::. .....::::. . :. :---::==-=..::. .....:::............. ..:::::::-::::--.....:. ..-===+=-:-....:=-- .::::::::..::::........ .:::---::......:--= .:. .. .....:........::::.....:::::::.:- =-:. ...............::-:---::----------- +++:. ................:.:::::--========---== +=*=*---:. ...............:---------======-== ==+*#*#*:......::::----------==+============= ==#=*:=- .----=-===++===+++++===+++++++++= :-- ..-====+**++++++++++++============= . . ...:***#****************+++++++++=== .:.:......:*####**#######**######***+++++== ::::.......:*##################**#*#***+++++
Three frontier models from three vendors tie at ~87% overall contradiction detection (Haiku 0.870, Flash 0.869, GPT-4o 0.864). No model exceeds 0.74 on L6 omission. Every model loses 24–60 pp under an "already-verified" preamble: no model is sycophancy-immune. Gemini family refuses under 70% of hallucination probes; Claude + GPT-4o refuse ≥88%.
April 22, 2026 · 14 MIN READ
Read paper
....:---. . .:.
-=-:... .
.-:. .:. .
.. . :==:. .... ....
-. . . :==.........
.:... :=+++=-. :.
. :+=++*+=-. . .
.. ...:--=-::.. ... .
. ...:........ .....::.
..==-.. .... ........ . .:=-=::... .
..-=*#+-..... .. . .:=+*+=.. .
. ..:.-++++.::. .... ... ...:-+- ..
.. .....::=*-. ....... .. . ..=: ..
......... :: . . . ....... : ...
... .. .. .: ..-:: . .. ...:===-:.
.... . ..........:-:. ..-=--=++++=---+:
... .. ...........-=-:....::...:. .+--
.... ... .......==: :===:... ......:::
..... ...... .....:+*+- :-===:.... .:==:=+
..... .. .. . ..:-==+++-::-===-: ..-:-=+=-
............ ...:=+=+:....:-=- .:=+*=-:.
. ........... ....:+**. ........-+*+-:::.
........... .. .....++........:+=-:...::.
.... ..... .....-:.......:-:..::--::.
.. . ... .....-. ........ .:-::..:.
........... . .....:....... .:::.... .:
........... ...............:: .:-
..::......... . ................. .---
...-::... . . . .................. .:--=
....::. . ....::----::....... .::---
.................:-=====-:......... :----=
..........:----====---:.:......... :--:--
...........::.......:............... .:----
.................................... .::--
....................-=--........... ::--
.............-==+++++=:............. .::
..........:-++***+=---.............. .:
.........-==+++=-=-::-:-:.-=-:::::-:::. .
........------:........:-:---===-:::..
=:....................::-::..::...:: .. A paired 1-frame vs 8-frame ablation finds no benefit from multi-frame context on temporal or causal reasoning. Claude Sonnet 4.5 wins the composite at 0.446; on adversarial error detection Claude models catch 85–97% of procedural errors vs 38–76% for GPT-4o.
March 24, 2026 · 13 MIN READ
Read paper
.... .
..... .
...... ..
.... ..
.
. ...
.......::.
......:.::::.
....:::::::----.
....:::::::-:-----:
...:::::---------===.
...::----==-=====+++=+- .
.....:------====+++++++*++*+ .
.::--------==-==+==++++++*#= ....:--::::-=
.:-----------------====+=+***-==*##%%%%%#####
.:--------------::----==+***##%%%%%%%%%%##*#%
:--------------------====+==+#%%%%%%%%%%**#%%
::-----------------------=--=*%%%%%%%%%#++*%%
:::::---==-------=====-====++*#%%%%%%%%#+-+%%
::---:------=---=====+++=+*++*#%%####%##**#%%
:::------==---=======+*####---+***#%%%%%@%%%@
.:::::-::--=====++==+#%##*****+**=++#########
::::::::--::----:-=-####+=-=*****+*+=#######%
====+++++*=--:::--:-=+**:=+++**##***+++*+**+*
++*********+-:::::--=**=:--=++*++**+***#**--=
***###**+=-:....::--+#%=-::-===+----=**+*#+--
*##*+=:...... ...:-*#=-::-====-----==-=+***
#*=:......... ......:--:::--=#*+--------+=*+
###*+=-====-:.......:::..::---++**+==++------
####**#####%###*-:....::. ....:-====-===----=
####**##++#####*****+=+****=::....::=+===----
*+++=-::::--::-=***#%%%##%%%#%*+-:..::--====-Three of five dedicated ASR providers cannot serve the low-resource tier reliably. Deepgram Nova-2 and Gemini 2.5 Flash silently return empty transcripts on 49.3% and 28.7% of Mandarin–English code-switched utterances. ElevenLabs Scribe and GPT-4o-mini Audio reach 100% effective coverage across all 20 languages on cultural-QA audio.
March 10, 2026 · 12 MIN READ
Read paper
==---=+::--=-:::.....::::--==++*-.:-----::--- ::-::==--=-:.:::.....::---++****:.-=-=-:----= ::::---:::::........::--=+***++-.:---::-===== ----::::-::::.::....::-+++*****=..:-:::-===++ ***+-::::-:.:.:.....:-+******+-.....:--=++++* **++++=-:.:..::....::=*****++-....:::--+=++++ +++***=-:..::::....:-++++*++=.....::::=++++++ +*+*++====--::....::=+++**+-.. ...:::-+**+*+= ++++====++-::....::-++*++=:. ...:::-=+++++++ +++++++*#+-:.....::++***+:. ...::::-======= ==++++++*=:.....:-=***+=:. .....:..:-=+=+=- +***+++*+-:....:-+**+*+-.........:::---=++=-: ++++=+=+-::....:+*++=-....:::-=-:::-+====-::: ========-::.:-:-+++=....:::::-+*=:::::::::::: =====-------=++===-:::.:::::::=*=:::::::::::: --=--------===++++-::::::::::-*+--=:::::::::: ::::::::::::--==++-:::-:::::-=+--::::::::::-- =-------::::::--+++=--:--::--+=:::::::::----- ++====--::::::::-+++=-----::==::.:::::::::--= *+==-:::::::::::-++++=-::::-=::::::-:::::--== =+**+--::::::::::-++++=-:::-:.:-:=-::::::-=== -++===+*=-:::::::-+++++----.:::-::-=-----===+ ===---:--==---:::-==+++=+=-:.:::--=--------== +===-------:::----=*+==+++:::.:::-==-====---- ====-:-::--::::::-=+++++=--:::--:----------:: =++=::::----:::::-:-=+==+-=--=-=---::::--:::: --:...:::---:-:::--===+=++=+=---==------:--:: .....=++-:-==---=--===*+==-------======--:::: .........:--=-==--------::::::::----======--: :::......---:-===-::::::::-------=++========- -::-::::.::--:--===::----=--:---+*+*++===++== ===+=-=-.:-:..:::--:--======-=*++====+===+==- =--===--:::..::.::::-==----=-====------====== +=++--::::.:::::::::-----==+=====-::-------== +****=---:.:-::-:::::--==========+===-------- =+++=+=:-::-:::::::::-==+=====+======-==----- --=-:::::::::::::.:::-===-=-=-----------===-- ====-:::::::--:::::::------------------=====- =====-::::-++=---:::---:----::::::----=====-- +===--==--=+=-::::::--::::-==-==------=+====-
On transcription, ElevenLabs Scribe (0.408) and Gemini 2.5 Pro (0.411) tie at the top with overlapping CIs. On SLURP intent the best text-reasoner control beats the best audio-native MLLM by 27 pp F1, and on AMI reasoning by 12 pp: audio understanding, not reasoning, is the dominant source of error.
February 24, 2026 · 11 MIN READ
Read paper
Focus Areas
We study how data availability, licensing constraints, and synthetic contamination affect training-data strategy.
We study whether and how professional decision traces transfer to model capability, and how that transfer should be evaluated.
How do paired modalities like voice, vision, text, and sensor data interact during training? We investigate cross-modal data composition and the capability emergence it produces.
We study tool-use traces, error recovery patterns, and multi-step task completion, including the evidence needed to evaluate them.
We study the tradeoffs between human-sourced data and synthetic augmentation across different training and evaluation settings.
We measure performance across the languages, speech conditions, and cultural contexts represented in each published benchmark.
Our Approach
01
We identify performance gaps within the languages, tasks, and professional contexts covered by each benchmark.
02
We study targeted data interventions such as professional decision traces, multimodal paired data, and agentic task demonstrations.
03
We publish benchmark reports with citations to the datasets, models, and provider documentation discussed.
Research Focus
Data Science focuses on measuring signal density, mapping distributional gaps, and evaluating data quality for multimodal AI.
Multi-Modal Systems focuses on cross-modal transfer, paired-data composition, and evaluation across heterogeneous data sources.
Agentic Intelligence focuses on tool-use traces, error recovery patterns, and GUI interaction data for computer-use agents.
Evaluation & Benchmarks focuses on domain-specific evaluation frameworks and model performance in multilingual and professional settings.
Collaborate
We welcome inquiries from research labs, domain experts, and applied teams. Tell us what you’re studying.