MindLens·Lab

In progress · updating weekly

Two preprints out.

The first 500 readings produced two papers, both openly available:

Collection continues toward the definitive Phase 1 paper.

Section 1 · Data accumulation

Two preprints deposited. Data continues.

The first target of 500 responses was cleared, and both Paper 1 and Paper 2 are now deposited on Zenodo (2 Sep 2026). Weekly totals below, cumulative, as the dataset keeps growing toward the definitive Phase 1 paper.

Phase 1 · Part 1 locked · 500+ responses (opening → 14 Aug 2026)

Two preprints on Zenodo (2 Sep 2026)

Phase 1 · Part 2 in collection · Sep 2026 onward

504 / 500readings

100%

Milestones
  • 1002026-06-01
  • 2002026-06-29
  • 3002026-07-27
  • 4002026-08-03
  • 5002026-08-10

Section 2 · Paper readiness

A research program, not a single paper.

Phase 1 splits its pre-registered hypotheses across a sequence of papers. Each paper waits until its data has real weight. The three below sit in different places on the runway.

Paper 1

Plurality is the norm

Published on Zenodo · September 2, 2026

  • H1 · Plurality claim

  • H2 · Verbal dependency ↔ agreement

  • H3 · Social complexity ↔ variance

43 adult readers, 33 short social clips, 442 responses. Most clips do not draw a single agreed reading, and about one response in four explicitly reports more than one emotion — even at the highest confidence.

Read on Zenodo →

Paper 2

Where the AI diverges

Published on Zenodo · September 2, 2026

  • H5 · AI ↔ human divergence patterns

Gemini 2.5 Pro on the same 33 clips: top pick differs from the human modal on about 40% of clips, and when the AI hedges its listed emotions overlap the human distribution 84% of the time.

Read on Zenodo →

Paper 3

Full Phase 1

Definitive paper · Phase 1 · Part 2 dataset · 2027+

  • H4 · Cue mediation of emotion choice

    developing

  • H7 · Reader-style consistency

    near target

  • H8 · Self-rated difficulty validity

    collecting

  • H6 · Cross-cultural variance

    long horizon

Closes out Phase 1 with a larger, deeper corpus: how perceptual cues shape emotion choice, whether individual readers are consistent, whether self-rated difficulty tracks disagreement, and how readings differ across cultures. Draws on the Part 2 dataset (Sep 2026 onward), which randomizes within-wave clip order and gates parental-consent verification.

"Ready" means the data required for that hypothesis is in place; formal test unlocks and effect sizes will appear on the pre-registration page once v0.2 is committed.

Section 3 · Timeline

Where this fits in time.

  1. 2025

    Originating paper — Curieux Academic Journal

    Artificial Intelligence in Emotional Intelligence Training for Autism (Kim, 2025). The paper that named the dataset-first pivot.

    See the paper →
  2. 2026 · summer

    Phase 1 · Part 1 · data collection

    First-target threshold of 500 responses cleared on 13 Aug 2026. Analytical cutoff 14 Aug 2026; dataset then frozen as Phase 1 · Part 1 · v1.0.

  3. 2026 · Sep 2

    Preprint 1 published on Zenodo

    Paper 1 — Plural Emotion Readings in Short Social Video Clips: A Descriptive Study. DOI 10.5281/zenodo.22251335.

    Read on Zenodo →
  4. 2026 · Sep 2

    Preprint 2 published on Zenodo

    Paper 2 — Where One AI Diverges from Plural Human Readings on Short Social Video Clips. DOI 10.5281/zenodo.22253369.

    Read on Zenodo →
  5. 2026 · Sep

    Phase 1 · Part 2 · data collection opens

    Same protocol on the same 33 clips, with two methodological improvements: randomized within-wave clip order (removes a within-wave dropout confound observed in Part 1) and gated parental-consent verification at session entry.

  6. 2027 +

    Definitive Phase 1 paper

    Paper 3 draws on the Phase 1 · Part 2 dataset (N > 100 target): cue mediation, reader style, difficulty validity, and cross-cultural comparison, once the data supports each strand.

Section 4 · Data availability by hypothesis

Where each pre-registered hypothesis stands on data.

Factual data-collection state only. Effect sizes and formal test unlocks appear on the pre-registration page once the v0.2 amendment is committed.

  • H1

    Plurality claim

    100%data ready

    Every live clip has enough readings for a distribution.

  • H2

    Verbal dependency

    100%data ready

    Requires curator-coded verbal_dependency and enough responses per clip.

  • H3

    Social complexity

    100%data ready

    Requires curator-coded social_complexity and enough responses per clip.

  • H4

    Cue mediation

    0%developing

    Needs deeper per-clip N so cue-partitioned sub-samples become usable.

  • H5

    AI ↔ human divergence

    100%data ready

    Needs an approved AI annotation per clip, verbal_dependency coding, and per-clip response coverage.

  • H6

    Cross-cultural

    0%long horizon

    Requires substantial per-country participant volume — the tallest ask in Phase 1.

  • H7

    Reader style

    72%near target

    Individual-level: participants need at least six answered clips.

  • H8

    Self-rated difficulty

    22%developing

    Needs the four-level difficulty scale answered by enough participants; newly moved to the session-summary flow to boost coverage.

Section 5 · Open science

Everything on this project is auditable.

Pre-registered protocol (v0.1)

Hypotheses, thresholds, and analysis plan committed before substantive data collection.

Read the protocol →

Phase 1 · Part 1 dataset (v1.0) — Zenodo

The frozen snapshot behind the two 2026 preprints. Cutoff 14 Aug 2026. CC BY 4.0.

Open on Zenodo →

Live dataset — Part 2 collection

Every new response joins the running dataset. No behind-the-scenes selection.

Explore the data →

Analysis code — public

Reproducible analysis scripts released on GitHub alongside the two preprints (MIT license).

GitHub repository →

No conflicts of interest

MindLens Lab is a non-commercial research project. No funders, no product ties.

Section 6 · Research notes

Decisions and observations along the way.

A running log of the small research decisions this project makes as it grows. Kept open so the reasoning is on the record, not just the result.

Sep 2026

Why we froze Phase 1 · Part 1 and randomized within-wave clip order in Part 2

After the two 2026 preprints went out, a post-hoc look at the Part 1 response counts turned up a systematic pattern: clips presented later in a wave's fixed order picked up meaningfully fewer responses than clips at the start of the same wave. Across the three collection waves, the within-wave rank correlation between clip position and response count was consistently strong and negative (Spearman ρ = −0.78 to −0.92; p ≤ 0.005 in each wave; ρ = −0.43 pooled, p = 0.012).

Two things follow from that. First, the Part 1 dataset used in the two preprints has a real dropout confound: later-position clips carry less statistical weight than their earlier neighbours, and any comparison between clips on the tail of a wave and clips on the head of a wave is standing on uneven ground. Second, this is a fixable problem, but only for future data — the fair move is to lock the Part 1 dataset exactly as the papers used it and to change the mechanism for the next round.

What we changed: The Part 1 dataset is now version 1.0 as of 14 Aug 2026, published as a Zenodo deposit (doi.org/10.5281/zenodo.22265192). Phase 1 · Part 2, opening in September 2026, keeps the same 33 clips and the same response instrument, but (a) randomizes clip order per session within each wave and (b) gates parental-consent verification at session entry rather than as a post-hoc exclusion. A subsequent definitive Phase 1 paper will draw on the Part 2 corpus.

Contribute a reading

Every reader moves the curve.

The most direct way to help this project reach its next paper is to spend about ten minutes reading a set of short clips.

Start the survey →