A multi-graph from OpenAlex
We ingest works, concepts, authors, institutions, and funders — 12k works across 7.8k fields — and connect fields by how they co-occur across the literature.
A prediction is only worth funding if the method that made it can be checked. Here is the whole pipeline, and the score it earns when we hide the future and ask it to rediscover the past.
We ingest works, concepts, authors, institutions, and funders — 12k works across 7.8k fields — and connect fields by how they co-occur across the literature.
We score pairs of fields that live in different research communities on shared neighbors (Adamic–Adar), author bridges, and recent momentum. The highest-scoring non-edges are the collisions: connections that don't exist yet but the structure says should.
For the top collisions, Claude reads the evidence and drafts a short brief — thesis, why now, who is positioned, what to fund — always with an explicit block on what would disconfirm it. The model writes prose; it does not invent the ranking.
We rebuild the graph from scratch as it stood at a cutoff year — every edge, every weight, every community, every field in the universe drawn only from papers published by then — rank the collisions, and measure the ROC-AUC against the truth: the pairs that actually fused afterward. The comparison is preferential attachment, the “rich get richer” null model you get for free by betting on whatever is already prominent.
Read this part. The engine orders the whole candidate set well, but at this cutoff none of the 157 verified future fusions land in its top 50 — the popularity baseline places 0. Treat the shortlist as a research prompt, not a hit list.
The engine ranks real future fusions above random ones (0.78 AUC), but none of them are in its top 50 — the ordering carries signal, the shortlist does not yet.
The concept graph is rebuilt from scratch as of the cutoff year from (:Work)-[:MENTIONS]->(:Concept) — edge existence, edge weights, the community partition and the concept universe all use only papers published by then. Scored: the composite ranking the product ships (Adamic-Adar x momentum x prominence). Ground truth: field pairs with no co-occurrence by the cutoff that then co-occurred at least twice.