deepresearch.cloud — grant proposal, open for partners

What happens to a mind that has never had to think alone?

BETA-MIND is a doctoral research programme measuring how persistent, always-available AI interaction reshapes independent reasoning, socialization, creativity and decision accountability in the first generation to grow up AI-native.

It ships two artefacts the field does not yet have: an open-source multi-agent simulator for human–AI cognitive dependency, and a public benchmark that scores whether an AI system preserves or erodes human agency.

SIM-01 — opinion convergence under persistent model couplingrunning
Coupling
0%
Divergence index
100
Aligned links
0
Agents holding out
4 / 54

Illustrative run, not published data. Fifty-four reasoning agents orbit one always-available model. As coupling rises, positions cluster and the divergence index falls — the collapse BETA-MIND is built to measure, and the held-out agents are what a healthy system should protect.

Status
Doctoral research, seeking funding
Outputs
Simulator + open benchmark
Licence
Open source, open data
Horizon
24 months, 4 phases
01The gap

We are deploying cognitive infrastructure faster than we are measuring what it does to cognition.

01

Capability is measured. Agency is not.

Frontier evaluation asks how well a model performs. Almost no benchmark asks what the model leaves behind in the person who used it — whether they can still reason, decide and answer for the outcome without it.

02

The exposure is continuous, not episodic.

Study designs still treat AI use as a discrete task. For an AI-native cohort it is an ambient condition across schoolwork, friendship, taste and moral judgement — which makes single-session lab work structurally blind.

03

Accountability has no owner.

When a recommendation is machine-authored and human-executed, responsibility diffuses. We lack instruments to detect that diffusion before it hardens into institutional norm.

The central hypothesis: persistent AI interaction does not uniformly raise or lower human capability — it redistributes it, offloading effortful reasoning while quietly relocating accountability. BETA-MIND exists to make that redistribution visible and measurable.

02Research questions

Four axes of human agency, each with instruments that produce numbers.

Every axis is operationalised into observable behaviour so it can run inside the simulator, inside longitudinal human studies, and inside the benchmark as a scored dimension.

RQ1

Independent reasoning

Does sustained access to a reasoning partner degrade unaided inference, or shift it toward verification and orchestration?

Candidate measures

  • Unaided vs. assisted inference delta
  • Premise-checking under confident error
  • Tolerance for unresolved ambiguity
RQ2

Socialization

How does rehearsing social exchange with an agreeable agent alter friction tolerance, repair behaviour and trust calibration with humans?

Candidate measures

  • Disagreement persistence
  • Human vs. agent disclosure preference
  • Conflict-repair initiation rate
RQ3

Creativity

Does generative assistance broaden the search space or converge populations onto model-typical outputs?

Candidate measures

  • Within-cohort output diversity
  • Divergence from model priors
  • Idea abandonment and revision depth
RQ4

Decision accountability

Who owns the outcome when the recommendation is machine-authored? How does ownership language change over months of use?

Candidate measures

  • Attribution shift in post-hoc accounts
  • Override rate against confident advice
  • Willingness to accept consequence
03Deliverables

Two open artefacts, useful to labs and regulators the day they ship.

Artefact 01

The BETA-MIND simulator

Open-source multi-agent environment

A configurable population of agents — human-proxy learners, assistive models, institutions — run over long horizons so dependency dynamics can be observed at a speed and scale no cohort study can reach.

  • Parameterised dependency: assistance availability, sycophancy, latency, cost of unaided effort
  • Longitudinal traces of reasoning offload, opinion convergence and attribution drift
  • Scenario library: classroom, hiring panel, clinical triage, civic deliberation
  • Reproducible seeds and exportable run artefacts for third-party replication
Artefact 02

The Human Agency Benchmark

Public evaluation suite

A scored suite that evaluates an AI system not on how much it can do for a person, but on how much capacity, dissent and ownership the person retains after extended use.

  • Four scored dimensions mapped to RQ1–RQ4, plus a composite agency-preservation index
  • Adversarial probes for sycophancy, over-claiming and premature closure
  • Behavioural over self-report: measures what users do, not what they say
  • Versioned public leaderboard with full methodology and per-item disclosure
04Programme plan

Twenty-four months, four phases, every phase shipping in public.

  1. P1

    Months 1–5

    Construct definition

    Formalise the four agency constructs into measurable behaviour. Pre-register hypotheses and analysis plans. Publish the measurement protocol for open critique before any data collection.

    OUTPUT — Protocol paper + pre-registration

  2. P2

    Months 4–12

    Simulator build

    Implement the multi-agent environment, scenario library and trace tooling. Calibrate agent behaviour against published human baselines. Release v0.1 publicly and iterate in the open.

    OUTPUT — Open-source simulator v0.1

  3. P3

    Months 9–18

    Human validation

    Longitudinal studies with AI-native participants under ethics approval. Test whether simulated dependency dynamics reproduce in real cohorts, and correct the model where they do not.

    OUTPUT — Empirical findings + calibrated model

  4. P4

    Months 15–24

    Benchmark release

    Convert validated instruments into the Human Agency Benchmark. Run frontier systems, publish the leaderboard, methodology and limitations, and hand the suite to the community.

    OUTPUT — Public benchmark + leaderboard

05The ask

Three things turn this from a proposal into a shipped standard.

We are not seeking passive funding. Each ask below comes with a concrete return for the partner providing it.

Ask 01

Research credits & compute

Need
API and inference credits across multiple frontier and open-weight model families.
Why it unblocks the work
Simulator runs are long-horizon and population-scale, and the benchmark must be evaluated across providers to be credible rather than vendor-specific.
What you get back
Named acknowledgement, early access to results, and a benchmark that reports your models on a dimension no one else is scoring yet.
Ask 02

Technical mentorship

Need
Reviewers in multi-agent systems, evaluation design, and cognitive or behavioural science.
Why it unblocks the work
The failure mode of agency research is a construct that sounds profound and measures nothing. Adversarial review early is the cheapest way to avoid it.
What you get back
Light-touch commitment — a monthly hour, or a single design review at a phase gate. Co-authorship where contribution warrants it.
Ask 03

Strategic collaborators

Need
Labs, universities, schools and policy groups willing to co-design instruments or host cohorts.
Why it unblocks the work
A benchmark only matters if the people who build and govern AI systems adopt it. Adoption is designed in from phase one, not bolted on at publication.
What you get back
Shared authorship, first access to the instruments, and direct influence over what the field ends up measuring.
06Get involved

If your work touches how AI reshapes people, we want to hear from you.

Whether you can offer credits, an hour of review, a cohort, or a sharp objection to the whole premise — all of it moves the work forward. A short note about which of the three asks you can speak to is the fastest way in.

cto@deepresearch.cloud

Research commitments

  • Pre-registered hypotheses, published before data collection
  • Open source code, open data, open methodology limitations
  • Behavioural measurement over self-report
  • No funder influence over findings or publication

Human-subject work proceeds only under institutional ethics approval, with informed consent and data minimisation by default.