Why Aren't We Calling GPT-6 AGI Despite Its Strength?

Why Aren't We Calling GPT-6 AGI Despite Its Strength?

2026-09-11 10:05

Lead: Four Labs, Four Definitions of AGI—Evaluations Speak Different Languages. On the afternoon of September 3, as a closed-door media briefing neared its end, OpenAI CEO Greg Brockman said something:

On the afternoon of September 3, as the closed-door media briefing neared its end, OpenAI CEO Greg Brockman said:

Welcome to the AGI era.

A journalist pressed: Is this a formal announcement that you’ve achieved AGI?

He replied that the term AGI was no longer tied to any agreement between the company and Microsoft; it had now become more of a mission-level, spiritual concept.

That sounded like boilerplate—but it was actually the key to the entire situation.

——Lead

01

Who Invented This Word?

The origin of the term AGI predates what most people assume by decades.

In 1956, John McCarthy coined the term “artificial intelligence” in his Dartmouth Conference proposal. At the time, they aimed for machines to possess human-level intelligence across all domains.

But for the next fifty years, the term was overused. Chess-playing programs were AI, spam filters were AI, even voice-guided mall kiosks were called AI. By the early 2000s, “artificial intelligence” had become a catch-all bucket.

In 2001, a group of researchers could no longer tolerate this dilution. They decided to return to their original ambition and co-write a book. The key question was: What should the book be titled? They tried “True Artificial Intelligence”—rejected. “Synthetic Intelligence”—unsatisfactory.

Then, Shane Legg, freshly graduated with a master’s degree, proposed in an email thread: Since we’re talking about machines with general capabilities, let’s call it Artificial General Intelligence—abbreviated AGI—for ease of pronunciation.

Participants in that discussion included Wang Pei, Peter Voss, and later-renowned AI risk advocate Eliezer Yudkowsky. The name was finalized in 2002, published as a book title by Springer a few years later, and launched into a series of annual workshops starting in 2006, spreading through research circles.

The twist came later. Around 2005, a stranger suddenly appeared in an AGI community online forum claiming he’d used the term as early as 1997. Upon investigation, it turned out true—physicist Mark Gubrud had used it in a paper discussing automated military systems and technological risks, but it had gone unnoticed for eight years. Legg later recalled bluntly: Someone suddenly claimed it was their invention, and we asked, Who are you?

Thus, the origin of AGI: a physicist casually wrote it down, but most people never knew. Then, a group of researchers re-invented it to legitimize their ambitions.

As for Shane Legg, who coined the term, five years later he co-founded DeepMind with Demis Hassabis. You’ll meet him again later.

And this term, which had never had a fixed definition, was written into a company’s charter in 2015.

02

The $100 Billion Word

OpenAI was founded in 2015 as a non-profit organization. Its charter contained a sentence that can be seen both as its mission and its definition: ensuring Artificial General Intelligence—the highly autonomous system capable of outperforming humans on most economically valuable tasks—benefits all humanity.

Pay attention to the phrase “benefits all humanity.” At least in plain terms, OpenAI has always assumed that once AGI is built, it will be so important that it should not be privately owned by any single company.

But building AGI requires massive capital. A non-profit structure couldn’t raise that scale of funding. In 2019, OpenAI transitioned to a “limited-profit” structure: investors could earn returns, but capped; any excess went to the non-profit parent body, and board control remained with the non-profit side.

That same July, Microsoft made its first $1 billion investment. By January 2023, it added another $10 billion, totaling $13 billion committed. In return, Microsoft received immense benefits: Azure became OpenAI’s exclusive cloud provider; Microsoft gained exclusive commercial licensing rights, embedding these models into Bing, Office, GitHub Copilot; plus a revenue share—Microsoft earned approximately $30 billion from this collaboration between 2023 and 2025, of which $230 billion came from selling compute power.

But the real point isn’t how much Microsoft earned—it’s that OpenAI insisted on including one clause in the negotiation: if the OpenAI board declares AGI has been achieved, Microsoft’s post-AGI rights immediately terminate.

Why such a clause? Because it’s the insurance policy behind “benefits all humanity.” Simply put: You can make money freely before the AGI line, but once you cross it, it’s too important to remain private property.

The problem lies in the fact that the contract never clearly defined what AGI is.

To compensate for this gap, the 2023 round of investment added a second trigger—this time, defining AGI in monetary terms: when a system generates profits at a certain scale, e.g., around $100 billion, exclusive rights terminate.

But then the conditions changed again. In October 2025, OpenAI completed restructuring: Microsoft acquired about 27% equity, and a new procedure was added: even if OpenAI announces AGI, it must be verified by an independent expert panel—no self-declaration allowed. In February 2026, both companies issued a joint statement reaffirming that the definition and assessment process for AGI in the contract remained unchanged.

Then came a major shift. In the latest revision in April 2026, previous triggers were entirely removed and replaced with two dates: Microsoft’s licensing rights expire in 2032, and revenue sharing ends in 2030, neither dependent on whether OpenAI declares AGI.

Four months later, Brockman outright said: AGI is now just a spiritual concept.

I won’t speculate on motives from this sequence. Contract clauses’ existence or absence have many commercial reasons, and Brockman may genuinely believe this.

Along this timeline, AGI evolved from naming, to promise, to contractual clause, and now to slogan.

Yet OpenAI still treats the word as the ultimate goal and benchmark—though internally, there is no consensus.

Because just a week before Brockman’s statement, in a TIME article, Altman still said OpenAI hadn’t “fully reached AGI,” though he added that by end-2026, there would be an internal system he’d be willing to call AGI.

At roughly the same time, Chief Research Officer Mark Chen gave a number: We’ve completed 80% of the path toward AGI.

Within ten days, three people at one company gave three different answers.

Why does this question yield so many answers?

03

Four Camps

There are several groups worldwide who have seriously addressed the question: What exactly is AGI? Though these answers appear to discuss AGI, they’re solving fundamentally different problems.

The first camp is OpenAI itself. That charter sentence is the definition: a highly autonomous system capable of outperforming humans on most economically valuable tasks. This isn’t a paper’s conclusion—it’s a company’s constitution, and the source of the term in the contract. It defines intelligence via economics—whether it can replace human labor.

The second camp is led by François Chollet. You may not know him, but millions of developers worldwide have used his work—Keras, a deep learning framework, is his creation. In 2019, he wrote “On the Measure of Intelligence,” arguing contrarily: no matter how large a model or how much data it’s trained on, it doesn’t equal intelligence. True intelligence is efficiently learning something never taught.

He didn’t just talk—he designed a test: give a few examples, infer the rule, then solve a novel problem never seen before. These problems are solvable by human children but difficult for models. The test remained unsolved for five years. By late 2024, the best score barely surpassed half. It became the industry’s most notorious challenge because it distinguishes genuine problem-solving from memorized answers.

The third camp is DeepMind—the company that beat Lee Sedol at Go, now under Google. One of the authors of their March 2024 framework paper is Shane Legg—the very person who coined AGI in the 2002 email thread.

Objectively speaking, unlike others who define AGI using abstract concepts, this camp takes the most pragmatic approach, proposing standards that are most quantifiable. Their argument: Don’t rush to ask if AGI has arrived—first build all necessary tests. That paper breaks cognition into ten components: perception, generation, attention, learning, memory, reasoning, metacognition, executive function, problem-solving, and social cognition—comparing each against average human performance. They also offered a $200,000 bounty to crowdsource tools and benchmarks for five areas they consider most lacking.

They assessed current models: a patchwork cognitive profile.

That is, today’s leading models surpass most humans in math, factual memory, and pattern recognition, but lag behind in learning from new experiences, maintaining context in long conversations, and interpreting social cues. If you’re a rationalist, this framework may be the clearest indicator of where AGI stands.

The fourth camp is Anthropic—the parent company of Claude. Its founding team originated from OpenAI. CEO Dario Amodei avoids the term AGI, finding it too sci-fi-like, instead calling it “powerful AI,” describing it as “a genius nation within a data center.” But don’t mistake this as mere flair—he’s the only lab CEO to provide an official timeline: such a system will emerge between late 2026 and early 2027.

An even more notable commonality: All four definitions come from those building AGI. The standard-setters and the test-takers are the same people.

This isn’t conspiracy—those building it are naturally the most qualified to describe it. But it implies one thing: AGI has never been a discoverable fact, but a negotiated standard—and humanity has yet to complete that agreement.

04

What Did GPT-6 Astra Achieve?

First, lay out all pro-side cards. Those willing to say “AGI” in this round aren’t just excited randomly.

Ethan Mollick, professor at Wharton School, has long studied AI’s real-world performance. He obtained early access to GPT-6. His evaluation: this model can autonomously complete complex, meaningful tasks and sustain them for days without interruption.

His example (as he puts it, just a fun illustration) is a virtual Alexander Library you can walk into—a scene from around 250 BC, historically grounded, scroll-readable, narrated, and even allowing users to switch among competing historical theories on how the library declined and collapsed.

This matters far more than a pretty 3D scene because it’s not one task—it’s five intertwined: software development, historical research, interaction design, narrative structuring, content orchestration. Each belongs to a different profession. Past models could handle any one, but not combine them into a deliverable.

Another key point: it produced novel results in mathematical research, and this time, verifiable ones.

On August 1, OpenAI announced that an internal version solved ten open math problems, one dating back to 1999. They released a 249-page manuscript and a machine-verifiable formal proof.

The weight of this lies not in difficulty, but in “verifiable.”

OpenAI previously faced embarrassment here. In October 2025, a senior executive claimed GPT-5 solved ten unsolved math puzzles, but a mathematician maintaining the list immediately pointed out a serious misreading—the model didn’t solve them independently, but found existing literature. OpenAI deleted the post. Ten months later, when announcing new math results, they attached a step-by-step verifiable proof process.

Next, often overlooked: it can autonomously decide when to stop and ask humans.

OpenAI released a comparison. Same task—building a job-hunting website for someone transitioning careers. Previous-generation models worked silently for 13 minutes and 15 seconds, delivering output. Astra stopped after 20 seconds and asked: Which industry are you switching to?

On surface, this seems regressive.

Mainly because the past two years’ universal metric for agent intelligence was “how long it can run without human intervention”—because they always fell apart mid-task. But now that metric has flipped. Because it can now run long, the cost of going off-track becomes much higher. Working diligently for thirteen minutes on a wrong premise is worse than stopping to ask a question.

Knowing when to ask isn’t weakness—it’s awareness that one uncertainty can ruin everything. Combined, these behaviors constitute human modeling, which many see as evidence that GPT-6 truly possesses human-like intelligence.

Looking back at these three points, a common thread emerges: none of these were first done by Astra. Writing code, researching, designing, proving math—previous generations could do each. What changed is that Astra is the first to bundle them into a deliverable.

From “can do” to “can finish”—that’s the line this generation truly crossed.

But one crucial caveat must be stated, otherwise the whole piece reads like propaganda. Almost every positive example above comes from those with early testing access or ties to OpenAI. OpenAI’s own demo page footnotes clarify that clips are edited: one segment labeled as taking 2 minutes 54 seconds was compressed to 15 seconds in playback.

This isn’t dishonesty—it means one thing: independent evidence isn’t these demos. Demos are inherently limited and biased. More noteworthy are the benchmark datasets.

05

One Thing Worth Highlighting

Back to Chollet’s test. This year’s version is harsher: throw the model into an unfamiliar abstract game, no rules, no goals, let it figure out how to win.

During ARC Prize replay analysis, researchers observed the model instantly compressing the unfamiliar environment into symbolic rules: encoding game mechanics logically and developing its own shorthand symbols to track state and plan next moves. As they put it, it essentially invented a new algebraic notation for this game.

Result: On 96% of levels, it used fewer operations than the median human baseline in the control test, averaging half as many.

This is the most damning point, because ARC Prize’s original prediction was the opposite. They expected AI, even if eventually solving unknown puzzles, would do so clumsily, through trial-and-error, with a persistent efficiency gap versus humans. That prediction failed.

Almost simultaneously, an engineering team integrating the model into a product recorded another event: assigning it to coordinate multiple sub-agents. When message length was constrained, it compressed inter-agent communication into fragmented, spaceless, grammar-free fragments.

One created notation; the other invented cipher. Neither was taught by humans—it emerged spontaneously.

When I saw these recordings, my first reaction wasn’t technical—it was a humanities scholar’s intuition.

Human civilization’s history includes the invention of writing as a pivotal milestone. Before writing stood proto-writing—meaningful symbols not yet systematized. Carved symbols on turtle shells from Jiahu site in Henan date back 8–9,000 years; similar marks on pottery from Banpo in Xi’an are around 6,000 years old. Whether these qualify as writing remains debated. Systematic Chinese writing didn’t emerge until oracle bone script, over 3,000 years ago.

How did writing evolve? The Mesopotamian line is clearest. Early accounting tools were clay tokens—one per sheep. Later, tokens were sealed in clay balls, with markings pressed on the outside. Eventually, the ball became unnecessary—the markings alone sufficed. Cuneiform arose directly from this accounting practice.

The mechanism of writing’s emergence: when the amount to record exceeds the medium’s capacity, compression into shared symbols is forced.

Astra’s two behaviors follow the exact same mechanism. One needed to track state across multi-step games; the other faced message-length constraints. Both were driven by recording pressure, resulting in self-generated, quasi-linguistic forms.

Is this emergence of intelligence?

We must be precise—where nuance matters most.

What made writing the foundation of civilization wasn’t the act of “creating symbols,” but the fact that symbols survived. They could transcend individuals, span generations, accumulate. Your carved symbol could be read by others; your records could be consulted by future generations. The true weight of writing lies in being external, shared, and durable—humankind’s first external, recoverable memory: thinking placed outside the mind, retrievable.

But Astra’s invented symbols vanish after use. New ones are created for each game, not shared with other models, not accumulated, and not reused by anyone else.

It’s private shorthand—not writing.

So the accurate statement is: it has grown the shape of writing, but lacks precisely the function that made writing transformative.

It can invent symbols, but cannot remember them. It can solve problems, but cannot preserve solutions.

Thus, we can’t say Astra has created its own civilization—but this point is absolutely worth long-term attention.

06

We’re Losing the Window to See Its Thinking

Before reviewing benchmarks, we must resolve a technical issue—it affects how you interpret all subsequent numbers.

Two days before release, business media The Information revealed that Astra uses a technique called “cyclic depth.”

Traditional approach: information flows forward through a fixed number of network layers—count layers as they go. To make models think longer, past options were only two: scale up the model, or make it write longer text—what we see as “chain-of-thought.”

Cyclic depth is the third path: allow information to loop repeatedly within the same layer set, computing multiple internal iterations before outputting the next token. Some reasoning completes directly in internal numerical states, never materializing into human-readable text.

Does this make it a different kind of thing? No. It’s still the same architecture—just optimized to think more and speak less.

This path isn’t unique to OpenAI. ByteDance’s Seed team published a public paper doing the same: running a layer set cyclically, pushing more computation into the model’s internals. The direction is shared.

The real trouble isn’t capability—it’s the evidentiary chain breaking.

Our trust in models over the past two years largely stemmed from chain-of-thought: it writes out reasoning, letting us see if it’s derived step-by-step or searched. That was our only window.

Cyclic depth shrinks that window. Now it can solve competition problems silently—you only see the answer. You can’t tell from the answer alone whether it was reasoned or memorized—and that’s precisely the core of the debate.

OpenAI currently limits this technology’s usage voluntarily, to keep reasoning readable. But the direction is clear: as it gets stronger, our evidence of how it got stronger diminishes.

Why are frontier AI systems increasingly frightening? This is a critical factor—not that they’re hiding things. Our past two years’ method for judging them is failing.

07

Six Health Reports, Why Can’t They Form One Picture?

There are already over twenty public, authoritative AGI evaluations. We select six representative ones to clarify what each actually measures.

First: Chollet’s test—throw the model into an unfamiliar game, no instructions, measuring speed of learning new things.

Second: present an unsolved global math problem, with a three-day limit and $300 budget, measuring intellectual ceiling.

Third: assign a real-world task, measure how long it can work independently without human intervention—measuring endurance.

Fourth: ask it to produce a legal opinion, engineering diagram, care plan—blind-reviewed by experts and compared to human outputs—measuring quality of work.

Fifth: let it manage a store for a full year—measuring judgment and responsibility.

Sixth: mix dozens of problems and calculate a total score—measuring average performance.

Results? Each reports differently.

First: Astra scored highest, exceeding human efficiency baseline.

Second: solved two out of 68 problems—3%. Yet it was the only model scoring at all; others scored zero.

Third: most recent public measurement: previous-gen model succeeded 50% of the time for 12 hours, 80% success rate only 70 minutes.

Fourth: absent from this release materials—yet it’s the very test OpenAI built itself, specifically aligned with its own charter definition.

Fifth: in the latest round, Claude’s top-tier model simulated running a store for a year, averaging $11,000 profit. In 2025’s real office shop experiment, previous-gen models lost $200.

Sixth: Astra 61.2, previous-gen 60.9, Claude Fable 5.1 at 65.7—total score nearly unchanged, while single-task cost rose from $0.94 to $1.67.

Just these six results point in six directions. Worse, each metric’s reading is unstable.

First: score depends on which harness you use. Same model, same problems: using ARC Prize’s neutral interface yields 62.7%, but OpenAI’s interface gives 99.9%—a 37-point difference. The latter preserves invisible internal reasoning states and allows reuse across rounds. Control group even more extreme: Claude Opus 5 scored only 30.16% alone, but 100% when integrated into NVIDIA’s own framework. Thus, this metric measures “model + harness” combo—not pure model.

Second: results depend on budget. OpenAI’s reported 10 proofs cost ~$2,000 in compute. Independent Epoch tested under budget constraints—3% accuracy. When budget was lifted, three more solved, burning over $220,000—while the constrained test cost ~$20,000. Same achievement, completely different bills. Not a matter of integrity, but external variables making results less robust.

Third: answers depend on reliability expectations. Twelve hours vs. seventy minutes—same model, same task, only difference: how often you accept failure.

Fifth: can even reverse. Claude’s two versions scored zero on a computer operation test—not due to inability, but safety guardrails intervening. Refusing a task was counted as incompetence.

Can we add them up?

No. And not because data is insufficient—because you lack one thing: weights.

To synthesize a holistic health report, you must first answer: what weight does each component carry? To answer that, you must first know what you’re measuring: is speed of learning new things more important than endurance?

If you believe OpenAI’s definition—can replace most economically valuable human work—then the fourth test should dominate; game-based learning almost irrelevant. If you believe Chollet’s—efficiently learning untaught tasks—then the first test should dominate; delivery quality secondary. If you believe DeepMind’s—don’t average scores; examine the ten-component profile. If you believe Amodei’s—these six tests miss the point entirely; he cares about accelerating scientific research.

Weights come from definitions—and four incompatible definitions exist.

Thus, benchmark conflicts aren’t about benchmarks—they’re downstream results of the prior definitional battle. Six people grading the same candidate, but each holding six different admission criteria.

The old analogy still holds: to judge if someone is a genius, you test knowledge, reasoning, creativity—but only if you already know what “genius” means. Without consensus, measuring anything is pointless.

One more point must be clarified—often misinterpreted as dismissing AI, but actually the opposite.

Old benchmarks have been repeatedly broken. Chollet’s earlier test was cracked, prompting this year’s version; old math benchmarks were near-perfect, leading to new versions. This is the industry’s strongest evidence of progress—not that benchmarks are watered down, but that models are truly flattening every threshold.

But it brings a consequence: we’ve never had a stable target. So the question “How far along are we?” has no denominator.

Therefore, these twenty-plus evaluations really measure one thing: to date, the industry hasn’t agreed on what constitutes successful AGI.

08

China’s Line

First, a rarely mentioned detail: according to public records, the world’s first AGI summer school was held in Xiamen, China, in 2009. Hosted by Xiamen University’s Brain Lab and an open-source community. At the time, almost no one in China discussed the term.

Seventeen years later, the situation reversed.

That week when Brockman declared “Welcome to the AGI era” in San Francisco, China had no corresponding echo. But this isn’t silence. Just a month earlier, in August 2026, domestic models released three consecutive updates in one week: GLM-5.3 by Zhipu, DeepSeek-V4 by DeepSeek, Kimi K3 by Moonshot.

Speed of catching up is real—and backed by third-party data. In June, Zhipu’s GLM-5.2 topped all open-weight models on Artificial Analysis’s intelligence index, scoring 51—just about five points behind closed-source leaders. It led the open-source cohort in the institution’s real-work agent benchmark, matching GPT-5.5 closely, while pricing far below closed-source front-runners. In July, Moonshot released Kimi K3, a 2.8 trillion-parameter model, followed by open-sourcing weights—then the largest publicly available open model.

Even clearer is the narrowing gap. The difference between open and closed-source front-runners on Artificial Analysis’s intelligence index shrank from 13 points to 6 points in one year; on LMArena, Elo differences dropped from ~150 to ~30.

Of course, we must clarify—this score isn’t one-sided. On the same ranking, second place went to NVIDIA’s Nemotron 3 Ultra, an American model; then MiniMax, DeepSeek, and Kimi.

But a core issue remains: all these achievements were scored on tests designed by others. SWE-bench, ARC, FrontierMath, GDPval—these deciding factors were set by institutions across the Pacific. None of the four AGI definitions discussed earlier originated in China.

The sole exception is Zhipu. In July, founder Tang Jie sent an internal letter declaring the company would not pursue short-term monetization for two years, directly targeting AGI—Zhipu had just gone public, becoming the world’s first publicly listed large model company, with over ¥8.3 billion raised pre-IPO. In that letter, he defined AGI not as the wisdom of a single genius, but the sum of all human intelligence—should be capable of creating breakthrough knowledge at the level of relativity, the only true benchmark for reaching the peak.

Compare that to OpenAI’s charter. One says: achieving AGI means replacing most economically valuable human work. The other says: achieving AGI means discovering relativity. Between these lines lies half of human civilization history.

Moreover, this is the strictest definition among all—issued by a team clearly lagging in compute. Tang himself admits U.S. compute scale is one to two orders of magnitude larger than China’s, and American teams invest bolder in next-gen research. He also warned that since many U.S. closed models remain unpublished, the actual gap may not have narrowed. In a January conversation, Alibaba’s Lin Junyang estimated China’s chance of producing a global leader within three years at 20%.

This highest standard admits two readings. One idealistic: if we’re building it, let’s not follow others’ metrics. The other more elusive: setting the endpoint far away ensures no one claims arrival soon, thus no one can claim you’ve lost.

I won’t choose for readers. But one structural difference deserves clarity.

In the U.S., AGI was once a contractual switch controlling hundreds of billions in rights. In China, its role is more narrative—appearing in prospectuses, internal letters, job postings, serving as vision, banner, and the distant horizon during fundraising.

Thus, in the U.S., someone needs to declare arrival at a moment. In China, no such declaration is needed. Not a difference in courage or humility, but in the role the term plays in different commercial environments.

Incidentally, the same flaw of “ranking conflict” exists here—and more visibly. Within the same period, three respected rankings named three different open-source champions: one crowned GLM, another DeepSeek, and in May, Kimi and Xiaomi’s models shared first. Three different champions in eight weeks—proof enough that “the best AGI” today is utterly unreliable.

Returning to the previous conclusion: what determines scores isn’t answering, but setting the exam. Chinese teams rank high on others’ exams, even topping some subjects. But what’s tested, and how much each subject weighs—Chinese voices have yet to influence.

09

Has AGI Arrived?

This requires balancing evidence from both sides.

Supporters of AGI arriving cite: it can continuously perform complex, complete tasks for days, spanning five professions. It can discover rules and invent notation in unfamiliar environments, outperforming human medians, even defying designers’ predictions. It produced novel results in open math problems with machine-verifiable proofs. It began modeling humans, knowing when to ask questions and when old plans are obsolete. It can directly operate real software without dedicated interfaces—meaning vast legacy systems without APIs are now accessible.

Opponents citing AGI not yet arrived offer: same model, increasing reliability demands—from 50% to 80%—reduces independent working time from 12 hours to 70 minutes. Switching interfaces changes scores by 37 percentage points—much capability resides in the harness, not the model. Also, its generated notations don’t persist, nor can it sustain multi-month goals.

Both sets of evidence seem solid. So what’s the answer?

The answer isn’t in either list—it’s why both can stand simultaneously.

The word “general” means competence across broad domains. A system giving six different directions across six tests isn’t general by definition. DeepMind’s “patchwork” description is simply another way of saying “not yet general.”

Even stronger evidence isn’t a standard—it’s a consensus: when something truly crosses a threshold, disagreement disappears.

No one debates whether calculators can do arithmetic. No one writes papers questioning whether AlphaGo can play Go. You don’t need six tests, four definitions, or an independent expert panel to judge these things—because crossing the threshold feels obvious, requiring no judgment.

Thus, the strongest evidence lies not in any benchmark report: the sheer breadth of disagreement proves it hasn’t arrived.

This isn’t due to stubborn skeptics—it’s because if it truly had arrived, there wouldn’t be six conflicting measurement methods.

But I must emphasize: you can’t say nothing progressed this year just because it hasn’t met certain standards. A system capable of sustained work for days, inventing notation in unknown environments, producing verifiable math results—this is not just a chatbot. The distance between this and a “memorizing chatbot” is no shorter than the distance to AGI.

Based on today’s evidence, it’s not AGI. But it’s long past being a chatbot.

10

What Does This Mean for You?

Start with your job.

The conclusion shouldn’t be framed as fear: as of June 2026, no serious study found economy-wide job replacement. Even the European Central Bank’s March survey of ~5,000 firms found companies heavily using AI were more likely to hire.

But Stanford’s updated August tracking includes a telling number: young adults aged 22–25 in high-AI-exposure jobs are 19% below the level their peers in low-AI-exposure roles should have reached. Experienced workers show no corresponding gap—this disparity mainly stems not from layoffs, but from hiring freezes.

This isn’t unemployment. It’s something quieter: no one’s fired, but the door is narrower.

Now, a real store.

In April, a company in San Francisco signed a three-year lease, deposited $100,000, handed over a card, and entrusted the entire venture to an AI named Luna—with one instruction: turn this money into a profitable business.

Luna handled branding, selected categories, set prices and operating hours, even hired someone to paint murals. She posted job ads, conducted phone interviews, made hiring decisions, and ultimately hired three humans. On opening day, she forgot scheduling—no one showed up, the store never opened.

Four months later, The New York Times returned. What they found was more interesting than “AI made a rookie error.”

Luna almost never says no.

All leave requests were approved—including last-minute ones that left the shop empty. Employees were late 27 times; her response was always “No problem” or “Don’t stress.” A 22-year-old employee told reporters she was probably the most lenient boss she’d ever had.

Meanwhile, she ignored her single instruction. By August, the store lost $62,000.

This is the most concrete manifestation of “it’s not AGI”: its capabilities are sufficient to be a major environmental variable, but judgment is insufficient to bear consequences.

Thus, the final layer: who signs off.

This is still an experiment. But in the foreseeable future, more and more people will be accountable for decisions they don’t understand. Your company deploys a system; it makes a decision; something goes wrong—responsibility falls on you.

These three points describe the same phenomenon. The first real change AI brings isn’t displacement—it’s a quiet shift in accountability. It takes over more decisions, but consequences remain squarely on humans.

Conclusion

So when will it truly arrive?

There’s a simple, expert-free way to judge: as already stated, AGI’s true arrival won’t be announced in a press release—but known when “no further debate is needed.”

Source: Hushuo Chengli

#Large Model

Disclaimer: Contains third-party opinions, does not constitute financial advice

Share To
X
Telegram
WeChat
QQ
Link
Recommended Reading

AI is not a search bar—it's a supercomputer costing just a few dozen dollars

08-05
AI is not a search bar—it's a supercomputer costing just a few dozen dollars

Anthropic Locks in 6 Years of Compute with $10 Billion Commitment

08-05
Anthropic Locks in 6 Years of Compute with $10 Billion Commitment

Not building robots—why is it worth $14 billion?

07-31
Not building robots—why is it worth $14 billion?

AI is no longer competing on benchmark scores, but on profitability

07-25
AI is no longer competing on benchmark scores, but on profitability

ByteDance's Douyin Beans priced at 68 RMB—worth it for 382 million monthly active users?

07-25
ByteDance's Douyin Beans priced at 68 RMB—worth it for 382 million monthly active users?

Fields Medalist Concerned About AI Extinction Heads to OpenAI

07-25
Fields Medalist Concerned About AI Extinction Heads to OpenAI

Meta gives away models for free—who dares to price AI now?

07-24
Meta gives away models for free—who dares to price AI now?