Tech leaders are field‑testing an “Einstein test” for AI: can language models trained only on historical texts independently rediscover landmark theories like general relativity? Early experiments suggest today’s systems fall short. Researchers report that LLMs struggle with abductive leaps—the creative, causality‑inventing jumps that defined Einstein’s work—often producing plausible but incorrect theories, or relying on nudges and modern hints. Attempts to build “vintage” models expose practical hurdles, from poor data quality to information leakage beyond cutoff dates, while separate studies show stronger progress in math where stepwise verification is easier. The takeaway: current AI excels at prediction, pattern mining and tool‑building, but translating sparse anomalies into new physical laws remains elusive—tempering near‑term hopes for AGI‑grade discovery.
Related articles:
— Can AI follow in Einstein’s footsteps?
— Highly accurate protein structure prediction with AlphaFold




























