Feature Image

Setaleur Aplamda

Pushing the horizons of Ai to a new level

OUR MISSION

Pushing The Boundaries of AI For A Stronger Vision

More about us

How we drive impact

About the laboratory

We are tackling the biggest dilemma in Artificial Intelligence

Our team is working to counter cautious, narrow learning in artificial intelligence and push it towards bold, ambitious learning.

More about our research

The Unnamed Architecture: On the Epistemic Structure of AGI Forecasting

Why a date without a defended mechanism fails the ordinary test of a scientific prediction





A prediction without a mechanism is not a forecast, it is a mood dressed up as data

-Momen Ghazouani 
Founder CEO Of Setaleur 
_______________________________________________________________________

Picture someone telling you, with total confidence, that we'll arrive in three hours. You wouldn't just start the clock. Your first question would be about the vehicle, not the number, because three hours by plane and three hours on foot describe two entirely different universes of possibility. The number means nothing until the method is named.

This is almost exactly the trick we keep hearing in the conversation around Artificial General Intelligence, and it's the trick we want to name plainly in this piece. In 2026 we've heard a Google DeepMind co-founder describe humanity as standing in the foothills of a singularity, with AGI perhaps four or five years out. We've heard an Anthropic CEO talk about a country of geniuses in a datacenter arriving as early as this year. We've heard an OpenAI CEO commit to 2028 and tell today's college students they'll graduate into a transformed world. Each of these lands as a confident, almost technical forecast. Almost none of them comes with the one piece of information that would actually make it meaningful: which research path, specifically, the speaker is betting will get there.

We don't think that omission is an accident. We think it's the central intellectual weakness in the entire AGI timeline conversation, and once you see it, you can't unsee it in the next headline.

Why This Is Hard

Every AGI timeline is really two claims fused into a single sentence. The first claim is visible: a date, a range, a horizon. The second is invisible: a belief about which class of system, scaled language models, embodied world models, some hybrid, something not yet built, is even capable of reaching general intelligence in the first place.

- Momen Ghazouani 
Founder CEO Setaleur 

That second, hidden claim is doing almost all the real work. If you believe the right path is to keep scaling transformer-based language models and wrapping them in longer-horizon agent loops, a five-year horizon is at least a coherent position, however debatable. But if you believe, as a serious and credentialed group of researchers now argues, that a system trained only on text can never acquire the grounded, causal understanding of the physical world that general intelligence requires, the same five-year number becomes close to incoherent, because the architecture it actually depends on has barely left the laboratory.

These aren't small differences in emphasis. They're different bets about what intelligence is made of, and yet the timeline almost always gets stated as if it floats free of that bet, as a pure empirical extrapolation, like a weather forecast, rather than what it really is: a conditional statement with the condition quietly removed.

We didn't invent this structure. Pierre Duhem, and later Willard Van Orman Quine, argued that no scientific hypothesis is ever tested alone, it's always tested together with a cluster of auxiliary assumptions, and a failed prediction never tells you by itself which member of the cluster was wrong. An AGI timeline is a textbook case of exactly this. The date is the hypothesis under test; the architecture is the auxiliary assumption smuggled in beside it. When a five-year prediction quietly fails to land, nothing about that failure tells you whether the error was in the timeline or in the architecture it silently assumed. Naming the hidden premise, then, isn't a stylistic preference for precision. It's closer to the minimum condition for the prediction to be falsifiable at all, a point we come back to directly below.

The Idea

Once you see AGI timelines as compound claims, a date wrapped around an unstated architecture bet, the natural next question is where that hidden bet comes from, and it turns out it isn't new at all.

The argument now associated with Yann LeCun and the world-models camp is, in its philosophical bones, at least fifty years old, and recognizing that lineage changes how the current debate should be read. In 1972 Hubert Dreyfus published What Computers Can't Do, drawing on Heidegger and Merleau-Ponty to argue that genuine intelligence requires being-in-the-world, a form of embodied, skillful coping with an environment that no symbolic system could achieve by manipulating representations alone. Dreyfus wasn't objecting to a specific algorithm, he was objecting to a structural feature of an entire research program, on the grounds that intelligence divorced from a body hits a ceiling no amount of extra symbol-processing can lift. Swap transformer for symbolic AI and the argument is, point for point, LeCun's.

Rodney Brooks made a related but distinct case from inside behavior-based robotics. In Intelligence Without Representation (1991), he argued that robust intelligent behavior can emerge directly from tight sensorimotor loops with the world, with no internal symbolic model mediating between perception and action at all, a claim that complicates any simple text-versus-embodiment framing, since it suggests representation itself, not just the absence of a body, may be the deeper issue. Stevan Harnad's symbol grounding problem (1990) gives the sharpest, most testable version of what LeCun gestures at with looser language: how do symbols acquire meaning if they're connected only to other symbols, never causally anchored to the world they're supposed to describe? A language model, on this reading, is a closed loop of symbols defined only in terms of each other, exactly the condition Harnad specified as insufficient for genuine semantics. Behind all of it stands John Searle's Chinese Room argument and its distinction between syntax and semantics, between manipulating symbols by rule and actually understanding what they mean, a distinction that echoes in David Chalmers's split between the easy problems and the hard problem of consciousness, transposed here onto a narrower question about cognition rather than experience.

We don't cite any of this for decoration. It tells us the world-models position isn't a contrarian reaction to a 2026 product cycle, it's the latest instance of a research program with a fifty-year track record of raising the same structural objection to whatever the dominant paradigm happens to be. That track record cuts both ways: it means the objection is serious and durable, but it also means the scaling camp's implicit response, that this time is different because the systems are simply larger, isn't new either. Dreyfus's critics said much the same about earlier waves of symbolic AI and were, for a period, wrong to write the field off early. History offers no clean verdict for either side here. It offers only evidence that the fault line itself is real, structural, and considerably older than either camp's current funding round.

We can also watch this old argument get repriced in real time. The clearest evidence that this fault line is still live, not merely of antiquarian interest, is that a Turing Award laureate put well over a billion dollars behind it in 2025. Yann LeCun spent twelve years as Meta's chief AI scientist before leaving that year to found AMI Labs on what has been described as the largest seed round ever closed by a European company. His argument, made consistently in public venues including an NVIDIA GTC keynote, is that language models are trained to predict text, not to understand or simulate the physical world, and that this is a structural ceiling rather than a temporary limitation. A system trained entirely on descriptions of gravity can produce fluent sentences about a falling glass without possessing anything resembling the intuitive physical model a toddler builds just by dropping things. The architecture he's betting on instead is the world model: a system trained on sensory and physical interaction data to predict how an environment changes over time, building an internal sense of causality, space, and object permanence from observation rather than from text about observation. His own research program, the Joint Embedding Predictive Architecture, is one instance of that bet; DeepMind's Genie 3, which simulates real-time three-dimensional environments, is another data point pointing the same direction, and Fei-Fei Li's push toward spatial intelligence belongs to the same family.

None of this is settled, and we don't think it should be treated as settled. At ICML 2026 in Seoul, AMI Labs' chief research officer argued in a keynote that language models understand the physical world only indirectly, through the secondhand filter of human text. At the same conference, Anthropic presented findings on what it called J-space, a self-organized reasoning structure it says emerged spontaneously inside Claude, which some read as evidence that scaled language models may be building internal world-like representations after all, without anyone designing them to. The debate has stopped being purely philosophical. It's starting to generate competing empirical evidence, and the evidence isn't one-sided yet.

Theoretical Contributions and Implications

We aren't presenting a benchmark result here, and we want to be precise about that rather than dress this up as more than it is. What we're offering is a diagnostic framework, and the clearest way to state its contribution is through Karl Popper's distinction between a forecast and a prophecy. Popper's Poverty of Historicism was aimed at a different target, the grand, law-governed predictions of history associated with Marxism, but its central complaint transfers almost without modification to AGI timelines. Popper's objection wasn't that historicist predictions were false, it was that they were unfalsifiable by construction, because they specified an outcome and a rough date without specifying a mechanism whose failure to operate would count as evidence against the prediction. A historicist prophecy can absorb almost any disconfirming event by reinterpreting it as a delay or a detour, rather than as a refutation.

Read carefully, AGI timelines have exactly this structure, and this is the specific claim our framework lets us make precisely. When Shane Legg's 2028 estimate or Ilya Sutskever's 2019 to 2021 estimate quietly slid, in each case by years and eventually by decades, that slippage was rarely treated by anyone, including the forecasters themselves, as evidence against the underlying paradigm. It got absorbed as a revision of the date, with the architecture left untouched. That's close to the exact move Popper flagged as the signature of pseudo-scientific prediction: a forecast that can be wrong about nearly everything except the one claim it was actually built to protect. Laid side by side, the current spread of expert predictions doesn't read like a maturing consensus narrowing toward an answer. It reads like a set of independent bets on which architecture will work, each one translated into a date as though the architecture question had already been settled, and each one defended, when it's defended at all, from inside its own assumptions rather than against its rival's strongest case. Demis Hassabis has moved his own goalposts multiple times within a single year, from a 2030 to 2035 window down to within the next five years, and in a May 2026 conversation attributed his growing confidence to a sense that the industry had found the right technical path, pointing to current AI agents as evidence, which is as close as a major lab leader has come to naming an architecture, and even there the claim is asserted rather than defended against the specific objection a former colleague is making loudly and expensively. Dario Amodei's country-of-geniuses framing, Sam Altman's 2028 date, and Mustafa Suleyman's twelve-to-eighteen-month horizon for white-collar automation all rest on the same unstated premise, stated with the same confidence and the same silence about the alternative.

We think there's also a second, complementary theoretical point here, drawn from Imre Lakatos's account of how competing research programs actually behave. Each program has a hard core of assumptions its adherents won't abandon, for the world-model camp, that grounded causal understanding can't arise from text alone; for the scaling camp, that sufficient scale and the right training objective are, in principle, enough, surrounded by a protective belt of auxiliary hypotheses that gets adjusted to absorb anomalies without touching the core. J-space is a natural anomaly for the world-model camp, and it's being absorbed by redescribing it as a functional analogue rather than genuine world-modeling. Genie 3's fluency at physical simulation is a natural anomaly for the scaling-skeptic camp, and it's being absorbed by classifying it as a different, non-linguistic architecture that actually proves the world-model point rather than refuting it. Neither move is intellectually dishonest, it's close to exactly what Lakatos predicted research programs do, and it's precisely why a single piece of evidence, however striking, was never going to settle this on its own.

What this changes, practically, is what counts as progress in the argument. It gives us a concrete way to separate the architecture claim from the timeline claim every time a prediction is made, and to ask each predictor to name and defend the mechanism their date depends on. It also points toward the kind of evidence that would actually move the needle: concrete, falsifiable tasks that discriminate between the two camps' core claims, for instance tasks that require inferring a physical outcome that isn't statistically recoverable from text alone, of the sort LeCun gestures toward when he notes that a system can describe a dropped glass shattering without ever having derived that outcome from anything but prior text describing similar events. If scaled language models keep closing that gap, that's real evidence against the world-model camp's central premise, not against a caricature of it. Findings like Anthropic's reported internal reasoning structures are exactly the kind of evidence that should shift these odds, in either direction, and should be tracked as such rather than filed away as a talking point for one side.

Impact and What's Next

It would be too simple to call these dishonest claims. It's more accurate to call them motivated ones, in the ordinary sense that anyone with a large financial or institutional stake in an approach has every incentive to speak about its timeline with more confidence than the underlying science supports. LeCun's billion-dollar bet requires the market to believe language-model scaling is a dead end; Amodei's and Altman's institutional futures require the opposite. That doesn't make either side's technical argument wrong, incentives aren't refutations, but it does mean a timeline stated by someone with an equity stake in one architecture should be read as advocacy wearing the clothing of forecasting, not as a neutral extrapolation from data.

There's a second, less individually psychological reason these forecasts keep the same shape, and we think it matters as much as the incentive story. Work on the sociology of expectations, from researchers like Harro van Lente, Nik Brown, and Mike Michael, argues that technological predictions of this kind are rarely neutral descriptions of a future state of affairs, they're speech acts that participate in bringing about the future they describe, by coordinating investment, recruiting talent, and setting the terms on which a technology will later be judged to have succeeded or failed. A five-year AGI timeline stated by a lab CEO is not only a claim about the world, it's an instrument for shaping the world it claims to describe, telling investors what horizon to underwrite, telling researchers what problem is worth a career, telling competitors what pace they need to match. A precise, falsifiable, architecture-specific prediction would actually serve that coordinating purpose worse than a confident but unmoored one, precisely because it could be checked. Journalists and audiences reliably ask when, and reliably fail to ask by what mechanism, and how do you answer the strongest objection to that mechanism. The timeline question produces a quotable, headline-ready answer; the architecture question requires the speaker to engage a genuine unresolved dispute and defend a side against a well-funded critic. Only one of those questions is being asked in most interviews, and it's the less informative one.

None of this means the timeline question is unanswerable in principle, only that it's currently being answered in a form built to resist falsification. Getting past that, in our view, takes three things. First, separate the architecture claim from the timeline claim explicitly, every time, and ask each predictor to name and defend the mechanism their date depends on. Second, keep pushing on the concrete, falsifiable tasks that could discriminate between the camps, and treat findings on either side, whether that's a world model closing a physical-reasoning gap or a language model developing something like J-space, as genuine evidence rather than a talking point to be absorbed. Third, take seriously the possibility that both research programs are, in Lakatos's sense, partially right, and that the more interesting question isn't which one wins but what a working handoff between them looks like, language and abstract reasoning on one layer, grounded physical simulation on another, each compensating for what the other lacks. Some of the field's most technically literate observers already treat that convergence, rather than a winner-take-all outcome, as the most likely resolution, and it's the direction we find ourselves watching most closely.

There's a harder question sitting underneath all of this, and it belongs to a different corner of epistemology entirely: what should a non-expert believe when genuine epistemic peers, researchers of comparable technical standing and comparable access to the evidence, disagree this sharply and this persistently. The literature on peer disagreement, particularly Adam Elga's and David Christensen's work on how a rational agent should update in the face of a recognized epistemic equal who disagrees, suggests the defensible response is rarely to just pick a side by force of personality or institutional prestige, and isn't obviously to average the positions into a bland middle estimate either. The more defensible response is to lower your confidence in any specific date while raising your attention to what the disagreement is actually about, which is exactly the architecture question we've tried to keep in view throughout this piece. Persistent, informed disagreement among genuine peers isn't evidence that the question is unanswerable, it's evidence that it hasn't been answered yet, and that anyone outside the disagreement should calibrate their confidence accordingly.

Closing

The next time a prominent figure states a year for AGI's arrival, we'd suggest the useful response isn't to argue about the year. It's to ask the prior question: what architecture is this number assuming, and has that architecture's central technical objection been answered, or just ignored. Until that question gets asked and answered, we think timeline predictions in this field deserve the same treatment as we'll be there in three hours from someone who hasn't yet said whether they're driving, flying, or walking. The number may still turn out to be right. Right now, though, it isn't a forecast. It's a preference, quietly wearing a forecast's clothes.

____________________________________________________________________________________________

View & Download PDF 

The Paper: The Vehicle Problem: Why Every AGI Timeline Is Secretly a Bet on Architecture

For citation: 
Momen Ghazouani. (2026). The Vehicle Problem: Why Every AGI Timeline Is Secretly a Bet on Architecture (Version 1). Zenodo. https://doi.org/10.5281/zenodo.21415631

BibTex: 

@misc{momen_ghazouani_2026_21415631,
  author = {Momen Ghazouani},
  title = {The Vehicle Problem: Why Every AGI Timeline Is
                   Secretly a Bet on Architecture
                  },
  month = jul,
  year = 2026,
  publisher = {Zenodo},
  version = 1,
  doi = {10.5281/zenodo.21415631},
  url = {https://doi.org/10.5281/zenodo.21415631},
} 

Post a Comment