Rich Bellantoni ·

Astra Is Extraordinary. That Doesn't Make Humans Obsolete.

What AGI and ASI actually mean, what Astra changes about autonomous work, and why capability, continual learning, and consciousness are different questions.


I’ve been using OpenAI’s Astra to work on a Unity game I play with friends and family, and the improvement over earlier models has been enormous. Especially the testing. I can have it walk through the game and play it, exercising the whole experience instead of just checking isolated pieces of code. It’s catching things no previous model caught for me.

Then I do a human playthrough, and there’s still a bunch of stuff I catch. Things a person would get hung up on, or things that don’t really look right even though they’re functional and technically work. Astra usually fixes them quickly once I point them out. But it keeps missing the human perspective, and the people who will play this game are human.

Both parts of that experience matter, because the capability can be incredible while the gap remains very real.

OpenAI’s Astra documentation describes multistep work across software, browsers, and code, with computer use and multi-agent orchestration. I think this generation is going to change what doing work on a computer means. I also think we’re getting careless about what that progress proves.

The analogy I keep coming back to is a bird and a plane. We built something that flies extraordinarily well without reproducing most of what makes a bird a bird. Likewise, an agent doing professional work doesn’t establish that we’ve reproduced human intelligence, solved learning, and put consciousness on a release schedule.

Those are different achievements. We should talk about them that way.

Before We Declare AGI, Let’s Agree on What We’re Declaring

Artificial general intelligence, or AGI, generally means broad competence across cognitive tasks, rather than excellence at one narrow task. The disagreement is about how broad, how capable, how reliable, and under what conditions. Matching an average person across many tasks is a different threshold from matching the strongest specialist in every field.

Artificial superintelligence, or ASI, is the stronger claim: intelligence that substantially exceeds human capabilities across a broad range of domains. Beating us at one thing, even something difficult, doesn’t establish that breadth.

Neither definition automatically includes consciousness. In their Levels of AGI paper, DeepMind researchers distinguish performance, generality, and autonomy, and explicitly separate capability from questions about subjective experience. You don’t have to reproduce the processes of a human brain to qualify under a capability-based definition.

That’s a reasonable approach. If a system meets a clearly stated AGI threshold, I don’t want to move the goalposts because the result makes me uncomfortable. But I do want the threshold stated before the victory announcement.

Autonomy describes how much work a system can carry out without intervention. Consciousness concerns whether there is subjective experience. An agent completing a spreadsheet is evidence about the first question. It doesn’t answer the second, however impressive the spreadsheet happens to be.

The Workstation Is Becoming a Workplace for Agents

Consider the business equivalent of that game-testing workflow: an agent moving between applications to produce a researched brief, a cleaned workbook, and a presentation. That’s the direction of travel I’m taking seriously with Astra and the models following it.

Add agents that divide the work, check intermediate results, and handle different tools, and the unit of automation starts getting larger. You’re delegating a workflow that previously needed someone sitting at a workstation, moving information around and deciding what to do next.

I expect that to become ordinary inside companies. How quickly, and in which roles, will depend on reliability, cost, access, and how clearly the work can be evaluated. Calling autonomous agents the norm everywhere already would get ahead of the evidence. Planning as though they’re a passing novelty would be a mistake in the other direction.

Some tasks will disappear into that automation, and some jobs will change or disappear with them. A system doesn’t need consciousness to create economic disruption.

But a swarm doesn’t automatically solve judgment. Several agents can check different things, or they can reinforce the same mistaken assumption. If every agent accepts the wrong definition of revenue, dividing the analysis among ten of them can produce a very polished answer to the wrong question.

In my game, functional success can still leave a human player struggling. In a business, a technically correct artifact can still miss what the people using it actually needed.

Remembering an Instruction Isn’t the Whole Problem of Learning

“Self-learning” is another phrase doing too much work.

AI already learns without a human labeling every example. Self-supervised training, reinforcement learning, and other approaches make that an established part of the field. The question isn’t whether machines can learn. Clearly they can.

The distinction I’m interested in is between adapting during a task and reliably accumulating transferable knowledge through experience.

An agent can read an error message, change its approach, and save a useful note. That’s valuable adaptation. It doesn’t, by itself, establish that the underlying model has acquired a durable new skill. Updating the information a system retrieves and updating its learned parameters are different mechanisms; OpenAI’s optimization guidance similarly distinguishes prompting from further training.

The practical test is what happens later. Does the system apply the lesson to a different problem? Can it recognize when the lesson no longer holds? Can it keep learning without damaging capabilities it already had?

People fail those tests too. Anyone who has managed an organization knows that experience doesn’t automatically become wisdom. But we can learn from a conversation, a failed experiment, or an unfamiliar situation and carry that forward into a very different part of life.

That’s the combination I remain skeptical we’re close to reproducing robustly across open-ended settings. More context, better memory, and longer autonomous runs are progress. They aren’t interchangeable evidence that the whole learning problem is solved.

There Are Still People Behind the Intelligence

The human contribution isn’t only the person checking the final answer. It runs through the creation and evaluation of training data as well.

Data annotation and labeling cover a range of work, from categorizing examples to writing demonstrations and assessing specialist answers. RLHF, reinforcement learning from human feedback, is one way human judgments influence training. These terms aren’t synonyms, and they don’t describe every part of a frontier model’s development.

But the work is real. DataAnnotation’s current listings and pay guidance advertise general projects starting around $25–$30-plus an hour, multilingual work from $20-plus, and coding, STEM, and professional projects at $50–$100-plus. Those are advertised rates from one platform, not a survey of what every worker earns. They do show companies purchasing specialist judgment, including from people with advanced degrees and professional credentials.

There’s something worth sitting with there. In the same industry discussing the replacement of expertise, there is a market for people explaining what a good answer looks like and why an apparently good answer is wrong.

That doesn’t mean each response requires a human behind a curtain, or that every conversation continuously updates the model. Training and deployment are different stages. It means the appearance of independence at the interface can hide a substantial human contribution upstream.

Human intelligence hasn’t vanished from the system just because it isn’t visible in the chat window.

Model Collapse Is a Warning About Losing Contact With the Source

The model-collapse paper by Ilia Shumailov and colleagues matters here. It examines what happens when successive generations of models learn from data produced by earlier generations. Under the conditions studied, errors compound and information about the original distribution is lost, including less common patterns.

What it doesn’t establish is a universal rule that every model collapses after two generations of synthetic data. The experiments span different models and training arrangements, with different degradation patterns. It also isn’t a demonstration that an individual model can only reason two steps beyond its training.

That distinction matters because later work, Is Model Collapse Inevitable?, found that accumulating synthetic data alongside the original real data avoided collapse in the settings tested. Replacing the original data and retaining it are meaningfully different training choices.

So the useful conclusion is about how we preserve information and evaluate what we’re adding. Synthetic data can help. A generated answer that has been checked against an experiment, a program execution, or a proof has a different evidential basis from another plausible paragraph accepted because it sounds right.

As AI-generated material circulates online, I worry about how easily the distinction between an observation and a recycled assertion can disappear. A claim repeated on a thousand pages doesn’t become a thousand independent observations.

For data leaders, that makes provenance and evaluation more valuable. We need to know where information came from, what checked it, and whether our pipelines preserve unusual cases rather than quietly sanding them away.

That’s familiar data work. Its importance doesn’t shrink because the downstream system is more impressive.

Creativity Is Real. Unlimited Self-Improvement Is a Further Claim.

I don’t think the best argument for human creativity is that a machine can never produce anything new. It can, and dismissing every useful result as remixing doesn’t get us very far. Human creativity draws on prior knowledge too.

There are stronger examples than a clever paragraph. Google’s AlphaEvolve combines language models with automated evaluation and evolutionary search to discover and improve algorithms. DeepMind reports applications to mathematics and to its own computing infrastructure, including AI training. That is meaningful invention, and it is evidence that AI can contribute to improving AI.

The important part is the surrounding process: proposals get tested, scored, and selected. The system has a way to distinguish an improvement from an attractive mistake.

In The Takeoff Story Has a Circular Problem, I questioned the leap from today’s models to systems that engineer their way past their own limitations. There’s a qualification worth making explicit: having limitations doesn’t prevent a system from helping remove some of them. Humans do that constantly.

What I still question is the assumption that these improvements will compound without encountering harder problems of evaluation, experimentation, or generalization. Automating an engineering task, improving a training kernel, and independently discovering a fundamentally better learning architecture are different achievements. Success at one doesn’t settle the others.

Nor does our incomplete understanding of the brain prove machines can’t get there. Neuroscience has already identified relationships between creative performance and interacting brain networks, including in Beaty and colleagues’ research on creative ability. There is knowledge here, even if it isn’t a complete engineering recipe for human creativity.

My concern is the missing demonstration: sustained discovery across unfamiliar problems, with reliable ways to recognize mistakes and learn from them. I won’t turn that concern into a claim of impossibility. I also won’t treat a forecast as though the demonstration has already happened.

The Bird-and-Plane Distinction Cuts Both Ways

Evolution didn’t have a plan to build us. Yet the depth of the result is easy to lose sight of when a model does something astonishing on a screen.

A human being learns while moving through a physical and social world, with a body, needs, relationships, and consequences. We bring those experiences into decisions about what matters and what to attempt. That combination deserves attention even when some of its individual tasks can be automated better than we perform them.

The flight analogy helps me hold both ideas at once. Birds helped inspire human flight, but aviation advanced through engineering that departed substantially from them. We didn’t need feathers to cross an ocean, and AI doesn’t necessarily need biological neurons to exceed us at difficult intellectual work.

Equally, a plane’s superiority at carrying passengers tells us little about a bird’s ability to find food, learn its environment, or raise its young. You have to specify the comparison.

That doesn’t guarantee anyone’s job. Aircraft displaced other ways of transporting people and goods; the economic consequences were real. But the achievement didn’t make biological flight meaningless, and a faster answer generator doesn’t make the rest of human life an obsolete implementation detail.

World Models Are a Direction Worth Watching

This is why Yann LeCun’s work on world models and joint-embedding predictive architectures interests me. It’s an attempt to learn useful structure from the world beyond the task of generating language.

Meta’s V-JEPA 2, for example, learns from video and predicts representations of what happens next. Meta demonstrated an action-conditioned version for robot planning after additional training on robot data. The underlying idea is to learn enough about how situations evolve to anticipate outcomes and plan actions.

I see promise there, particularly for connecting observation, prediction, and experience. But a world model isn’t automatically a conscious model, and JEPA isn’t a demonstrated guarantee of AGI. It may become part of a broader system alongside language models, memory, search, and tools. I wouldn’t assume the eventual answer has to belong to one architectural camp.

Consciousness deserves its own evidence. Research led by Patrick Butlin and Robert Long develops indicators informed by theories of consciousness, while acknowledging uncertainty and the possibility of future systems satisfying those indicators. That is a very different exercise from deciding that fluent conversation or professional competence proves subjective experience.

I remain unconvinced that these capabilities establish conscious AI. I also don’t think we have an honest countdown to it, in either direction.

We Can Adapt Without Writing Ourselves Out

There’s a lot of money and ambition attached to what people believe the next model represents. I want to keep asking what has actually been demonstrated, especially when the answer changes how organizations treat people.

Use the tools. Learn to delegate larger pieces of work, and measure whether the result holds up. The person who understands the business, can challenge the objective, and knows where to look for a hidden failure has more to contribute than manually producing every intermediate artifact.

But don’t assume those abilities maintain themselves. People need opportunities to practice, investigate, make decisions, and learn from the consequences. If we automate the work through which expertise develops, we need to think about how the next generation acquires it.

I wrote in The AI Story Is Breaking about protecting the people who know what good looks like. My Unity experience makes that concrete: I want Astra doing more of the testing, and I still want a human playthrough before friends and family sit down to play.

And there’s something beyond the business case that I don’t want to lose. Our worth isn’t a benchmark lead we have to defend against the next release. Relationships, curiosity, responsibility, and the choices we make about our lives don’t become worthless when a machine gets better at a task.

The technology can be fantastic. Humans can still be special. I don’t see a contradiction there.


We built the plane because we wanted to go further. We can do the same with AI without deciding the bird was a mistake.