A multi-graph from OpenAlex
We ingest works, concepts, authors, institutions, and funders — 8.6k works across 6.8k fields — and connect fields by how they co-occur across the literature.
A prediction is only worth funding if the method that made it can be checked. Here is the whole pipeline, and the score it earns when we hide the future and ask it to rediscover the past.
We ingest works, concepts, authors, institutions, and funders — 8.6k works across 6.8k fields — and connect fields by how they co-occur across the literature.
We score pairs of fields that live in different research communities on shared neighbors (Adamic–Adar), author bridges, and recent momentum. The highest-scoring non-edges are the collisions: connections that don't exist yet but the structure says should.
For the top collisions, Claude reads the evidence and drafts a short brief — thesis, why now, who is positioned, what to fund — always with an explicit block on what would disconfirm it. The model writes prose; it does not invent the ranking.
We hide everything after a cutoff year, rank collisions from the graph as it stood then, and measure the ROC-AUCagainst the truth — the pairs that actually fused afterward — versus preferential attachment, the “rich get richer” null model.
ROC-AUC of cross-community Adamic-Adar vs preferential-attachment; ground truth = concept pairs whose first co-occurrence fell after the cutoff