Google DeepMind reports that AlphaEvolve improved the best-known solutions to roughly a fifth of more than fifty open mathematical problems it tested. These are new, checkable results. They raise a question that a list of answers cannot settle: where, exactly, did the discovery happen?
AlphaEvolve proposes programs, tests them with automated evaluators, and uses the results to guide further search. People selected the problems and built the system that could recognize an improvement. The machine found answers people had missed within that arrangement. DeepMind's account of the system leaves a different question open: could a machine originate the question, the useful representation, and the way to test an answer?
That is the part of intelligence I most want to isolate. Can a machine enter an unfamiliar world, notice what needs explaining, and invent a science of its own? If it can, what would that tell us about the other question often attached to intelligence: whether a mind must be conscious to understand anything?
Give it a world it hasn't seen
Our world is a poor place to run this test cleanly. Current AI systems have learned from enormous amounts of human writing, including scientific explanations and the mathematics used to express them. When one explains a phenomenon we already understand, the source of its concepts is hard to trace. Human scientists also inherit a vast intellectual culture, so demanding that an AI start without any inherited ideas would apply a standard no person has met.
Here is a more useful experiment. Build several simulated worlds with coherent rules unlike the physics we know. The rules should generate many observable effects, allow instruments to be built, and permit theories that work at first but fail when measurements improve. The designers know the rules. The participants do not.
Give separate communities of people and AI systems comparable observations and actions. Both groups can manipulate their surroundings, store results, communicate, and build tools. Neither receives a list of scientific concepts to find or a score that says whether its latest theory is getting warmer. The initial instruction can be broad: learn enough about this place to predict and control what happens in it.
This would test a community rather than a lone genius. Discoveries accumulate through shared records, criticism, specialization, and better instruments. That is the modest meaning of "civilization" here. No one needs to build a city.
Watch the path to an explanation
A final leaderboard would lose much of the evidence. I would want to see which observations each group considered surprising, which measurements it invented, and which experiments it chose when two explanations fit the same data. Does its explanation predict a new effect? Does a concept developed in one setting help it understand another?
The strongest case would involve a theory that initially works. Imagine that one group explains a wide range of effects using a simple quantity. With better instruments it discovers small, reliable exceptions. It could keep adding special cases, or it could replace the original quantity with a new way of describing the world. That replacement matters if it predicts an effect in a part of the world no one has examined, and the effect appears when tested.
The new concept need not be a primitive in the simulator's code. Temperature, for example, can describe a population of particles even when the simulator assigns no "temperature" field to any one of them. A machine community might discover an equally useful higher-level quantity in a world whose designers never named it. We should judge that idea by what it explains, predicts, and lets the community do, rather than by how strange it sounds to human readers.
The test would need repeated worlds and careful budgets. Humans and machines run at different speeds; equal clock time would be a poor measure of opportunity. Researchers could track the observations, experiments, communication, tools, and computation each group used. They could also withhold some phenomena until after a theory is proposed, reducing the temptation to reward a story that merely fits everything already seen.
A result would still need interpretation
"From scratch" is an approximation. Humans arrive with perceptual habits shaped by evolution and experience. AI systems arrive with architectures designed by people and concepts learned through human language. A synthetic world reduces the influence of familiar facts; it cannot erase either group's starting assumptions. Even the simulator and its interface reflect decisions made by its designers.
This is why I would look for a pattern across many worlds, not a single dramatic result. If machines repeatedly made useful conceptual revisions without receiving the crucial concepts or tests from people, the claim that they can only rearrange inherited human ideas would become much harder to sustain. If they consistently optimized within an existing framework while human groups found better ones, that would be a real puzzle. It would take more work to show that memory, experimental access, motivation, or the interface was not responsible.
What about consciousness?
For this experiment, scientific understanding means proposing explanations, designing tests, revising concepts, and making predictions that survive contact with the world. Consciousness, in the sense relevant here, concerns subjective experience: whether there is something it is like to be the investigator. These questions are connected, but success on the first does not directly measure the second.
Suppose a machine civilization passes the test. We would have strong evidence that a machine system can develop explanatory science under those conditions. We would not know whether the machines involved had subjective experience. Their success therefore would not, by itself, prove that consciousness is unnecessary for scientific understanding.
Now suppose the human groups repeatedly succeed where well-equipped machine groups fail. That would identify a cognitive difference worth explaining. Consciousness might be part of an explanation, but the result would not single it out. We would still have to distinguish it from differences in architecture, learning history, embodiment, and the incentives built into the experiment.
The test would sharpen the argument about machine understanding. It could show what explanatory work these systems can do when given a world instead of a problem set. It cannot tell us whether there is experience behind that work. Any claim that consciousness is essential would then need to identify where, in the path from anomaly to successful theory, experience makes the difference.
Sources and open questions
- Google DeepMind's AlphaEvolve report describes the reported mathematical results and the system's language-model, evaluator, and evolutionary-search components.
- AI Feynman is an example of existing automated equation discovery from data. Its authors designed a physics-inspired symbolic-regression method and tested it against known equations; the proposed experiment asks how a community would choose its own questions and explanatory concepts.
- The proposed worlds, comparison, and interpretation are a thought experiment. The design needs a broader review of related work and a clearer account of how to make worlds unfamiliar to both groups without quietly supplying the concepts being tested.