SYNAPSE

← work

model evaluation infrastructureSYNAPSEErnst & YoungPLAN Labmodel evaluation infrastructureSYNAPSEErnst & YoungPLAN Lab
sep 2025 – presenturbana, il

SYNAPSE

PLAN Lab, advised by Dr. Ismini Lourentzou

Asking a single model whether a scientific claim is true is unreliable: it has no way to check its own work, and it tends to favor whatever evidence confirms the claim. We’re building a system that breaks a claim down into a dependency graph of subclaims, so that if a foundational piece fails verification it propagates up through the rest, and checks each piece against papers, data, and simulations with a permanent adversarial agent whose only job is to argue against the others. It outputs a feasibility score with the reasoning and citations behind it, starting in materials science and expanding from there. Still working out what the right division of labor looks like, and how to tell confident-and-correct apart from confident-and-wrong.

TLDR: instead of asking one model whether a scientific claim is true, several models with different jobs and different tools check it together, and one of them argues against the others.

Anvesha presenting SYNAPSE research
presenting SYNAPSE!