AI at Prova
Put the power to find out in more hands.
Our purpose is to democratize the ability to find out what works in social policy.
The next advance in how we help people could begin in a neighborhood program, a school, or a public agency. More of those ideas deserve the chance to be tested and developed.
AI opens a remarkable possibility: putting serious research within reach of many more people working for social progress. We’re building Prova to make that possibility real.
01 · The opening
More complex investigations are coming within reach.
Serious inquiry brings many forms of work together: reasoning through a question, pursuing evidence, designing a study, building its instruments, and checking what the findings establish.
AI’s growing ability to contribute across that work changes what people and institutions can undertake. Prova’s founding thesis is that expanding this capacity can change the organization of social research.
Capabilities that extend the work
Evidence available by August 2026Work through a difficult question.
Weigh arguments, examine assumptions, and connect the implications.
Conceptual Reasoning Index · points
2024 model
August 2026 result
Opus 5: 95% confidence interval ±2.1 points.
Pursue evidence through tools.
Search across sources, follow leads, and locate hard-to-find information.
BrowseComp · questions answered correctly
April 2025 result
July 2026 result
1,266 questions. Reported agent systems use different configurations.
Carry out a demanding analysis.
Work with spreadsheets, documents, and code to resolve a question.
AA-AnalystAgent · correct on all five attempts
Gemini 3.7 Flash · high reasoning
80 questions across 14 domains. August 13, 2026 result.
Bring the capabilities into one investigation.
- Frame the question
- Pursue evidence
- Design & build
- Analyze & check
Read the measures and sources
- Reasoning. The Conceptual Reasoning Index combines argument assessment, consistency of beliefs, and decision-theoretic reasoning. Both models are scored on the same index; GPT-4o’s score is retrospective. The August 12 release reports results through August 10. Exact scores are available in the results table. Index points are not a ratio scale of intelligence.
- Evidence gathering. BrowseComp tests difficult fact-finding with short, verifiable answers. Deep Research’s April 2025 result and Sol’s July 2026 result use different trained systems and configurations. The latter is the single-agent result, available by August. This is not a controlled estimate of model-only improvement.
- Analysis. AA-AnalystAgent uses tools for code execution, web fetching, and image viewing. A question counts as solved only when all five independent attempts are correct (pass⁵). The squares show the aggregate share, not individual question identities. See the benchmark methodology and August model evaluation.
Supporting evidence · METRHow complex can the work become?
From minutes to multi-hour problems.
METR measures the difficulty of a task by how long a human expert would need to complete it. Its horizon estimates show the task duration at which an AI agent is predicted to succeed at a given rate.
- GPT-4March 20234 min95% CI: 1.9 min–8 min
- GPT-4oMay 20247 min95% CI: 4 min–13 min
- Claude 3.7 SonnetFebruary 20251.0 hr95% CI: 33 min–1.7 hr
- o3April 20252.0 hr95% CI: 1.2 hr–3.2 hr
- GPT-5August 20253.4 hr95% CI: 1.9 hr–6.8 hr
- Claude Opus 4.6February 202612.0 hr95% CI: 5.3 hr–60.6 hr
At 50% predicted success, Claude Opus 4.6’s task horizon is about 12.0 hr. The confidence interval is 5.3 hr–60.6 hr.
The tasks primarily cover software engineering, machine learning, and cybersecurity. They are self-contained problems with clear success criteria; the results do not measure complete social-research projects.
Source series updated May 8, 2026. This is evidence available by August, with no August measurement added or extrapolated. METR’s methods · Published data
Build the capacity to use them together.
Prova brings these developing capabilities together with organized research knowledge and the digital material institutions already produce. The aim is to investigate more demanding questions, build what a study needs, and support learning as a program develops. Lower costs can widen access to that capacity.
02 · What we’re building
AI created the intellectual foundation.
Prova is an AI-native company. AI created the organized body of research knowledge and relationships at the heart of our proprietary systems, connecting methods with the people, resources, and conditions a study depends on.
We use AI with that structure throughout research, design, analysis, and implementation. The question, evidence, local conditions, and professional expertise give the work its direction.
Reasoning, research & software.
Models contribute language, reasoning, coding, and tool use.
Prova’s organized research knowledge.
Research put to work.
- Studies
- Findings
- Instruments
- Software
What is a compiled prior?
Our starting research knowledge, organized for use. AI created this body of methods, assumptions, and relationships. It remains open to examination, correction, and revision. The bands above illustrate its scope.
A company organized around that capability.
AI helped make Prova itself possible, expanding what its founder could undertake in developing the systems and building an organization around them. It continues to contribute to strategy, internal operations, software, and product development.
Research and product development can both strengthen the capability. A study can reveal a missing method; a tool can make that method practical to use. Prova is being built to connect these forms of work.
Change the question. See the research change.
Are participants in work six months later?
- Design
Plan a follow-up
Measure employment six months after the program.
- Build
Create the instruments
Develop a survey and a collection tool.
- Carry out
Reach participants
Run the follow-up and track missing responses.
- Learn
Understand the outcome
Describe employment among respondents and examine missing follow-up.
03 · An invitation to researchers
Making research knowledge a testable source of AI capability.
We are exploring how much of the capability required for serious social research can be built into an explicit, revisable body of knowledge and relationships, then carried across generations of AI models.
Maintaining this knowledge separately from the model makes the structure itself available for examination and improvement. For AI and ML researchers, that creates a concrete experimental agenda.
Research agenda · Proposed evaluation
Test what the structure contributes.
A strong task-prompted model
Case materials and a carefully developed task prompt.
The same knowledge, in prose
Matched facts and relational statements, presented as text.
Explicit relationships
The matched knowledge, organized into an explicit structure.
- Design quality
- Source fidelity
- Expert corrections
- Total effort
Does the organization of knowledge add value beyond access to it?
Match factual and relational content and coverage in B and C. A tests what adding research knowledge contributes.
Proposed experiments. No performance advantage or completed evaluation is asserted.
Another question concerns how the work preserves agreed analytical commitments when results or incentives make them inconvenient. We invite collaborators to help establish which improvements are real, which transfer, and what deserves to be built next.
Related research on compound AI systems and DSPy provides a broader setting for this work. The comparisons above are Prova’s proposed research agenda.
04 · The responsibility that comes with it
The same capabilities raise the stakes.
Greater reach, lower costs, and continuing learning bring real possibilities. Each also asks something of how the work is designed and judged. A mistake in a reusable research structure can travel into later studies and tools.
Select a capability to explore the discipline it requires.
Investigate more
Extend research, design, and implementation.
Pollution
Weak work at volume. Errors carried into later work.
Reach more programs
Bring serious inquiry within reach.
Dismissal
A low price mistaken for a low standard.
Keep learning
Make inquiry part of ordinary operations.
Surveillance & target chasing
People lose a say. The score replaces the goal.
Check sources, methods, and tools. Challenge assumptions before they travel into another study.
Institutions also need arrangements in which an unwelcome finding can lead to useful action.
More evidence has limited value if a program cannot change, a funder cannot reconsider, or people cannot challenge how they are represented. The ability to learn depends on the authority, incentives, and relationships around the work.
05 · The larger ambition
Knowledge and capability should accumulate.
A study can leave a better instrument, a useful method, or a finding another team can examine. Every investigation should give the next one a stronger start.
Advances in models and improvements in maintained research knowledge offer two ways for the capability to develop. Establishing what can travel between questions and settings is part of the work ahead.
↓ Can leave useful assets
↓ To inform further work
The longer horizon reaches across institutions.
We want useful knowledge to survive individual projects, leadership changes, and funding cycles. The ambition is a growing capacity for institutions to cooperate, test ideas, learn from unsuccessful attempts, and act on what they discover.
For capital committed to social progress, the question is what becomes possible when more institutions can sustain inquiry. For Prova, the task is to build a company whose research, systems, and products help bring that future within reach.
Build with Prova
Help develop the capacity to find out.
We welcome foundations and investors interested in developing this capability, researchers who want to test it, and institutions willing to shape its real-world use. The next stage is to establish quality, transfer, total cost, and the conditions for responsible adoption.