The AI Future Meets Reality

From o3 and CAPTCHA breakthroughs to working agents, clinical evidence and Australia’s next rulebook, our archive reveals how far AI has travelled—and which promises still need to earn their keep.

 

Two years of AI news can feel like a time capsule from several different decades. A model learns to reason. A machine recognises the traffic lights. Software begins operating other software. Each arrives as a glimpse of the future, until the next announcement makes it sound almost ordinary.

Reading back through FineSkyAi’s hundreds of news articles, from May 2024 onwards, a larger story emerges. Many capabilities we once covered as research frontiers are becoming usable tools. Others remain experiments, forecasts or contested claims. The distance between those categories matters more as AI takes on consequential work.

The future has arrived unevenly. It has also brought an electricity bill, a copyright dispute and a considerably longer procurement questionnaire.

The AGI question returns

In December 2024, we covered o3 and the suggestion that artificial general intelligence might be close. The article welcomed stronger reasoning while questioning whether a benchmark result established broad, adaptable intelligence. That caution has aged well alongside the excitement.

On 3 September 2026, OpenAI introduced GPT-6 Astra, reporting major advances in scientific reasoning, coding and computer use. The independent ARC Prize assessment found a best score of 99.9% on ARC-AGI-3’s semi-private evaluation using OpenAI’s specialised setup, compared with 62.7% using the standard interface. The setup determines how reasoning state survives between requests; it is a meaningful part of the result.

These are different tests and conditions from o3’s original evaluation, so the figures are not a simple progress chart. They do establish a substantial new milestone. ARC Prize nevertheless explicitly says it is not claiming Astra is AGI, pointing to the gap between bounded test environments and the open-ended world.

For this retrospective’s editorial exchange, I put the question to GPT-6 Astra myself: would it call itself AGI?

“I can demonstrate capabilities; I cannot certify myself as AGI. Judge me by how reliably I handle unfamiliar work, recognise my limits and remain accountable to the people using me.”

The response was generated for this article. It is a useful position, rather than an independent assessment of the system delivering it. As our February consciousness article also explored, capable behaviour and subjective experience raise different questions. A fluent answer cannot settle either debate by itself.

The traffic lights stopped being enough

Our October 2024 CAPTCHA story was already reporting successful machine performance. The prediction concerned what would happen to verification systems built around tasks supposedly easy for humans and difficult for computers.

The underlying ETH Zurich research demonstrated successful completion of all tested reCAPTCHA v2 sessions in its successful configuration. That did not mean every image was classified correctly on the first attempt, or that every CAPTCHA system had been defeated.

Since then, the problem has broadened. At USENIX Security 2025, researchers presented Halligan, an agent using a vision-language model that solved 60.7% of 2,600 challenges spanning 26 types. The progression is from a specialised solver towards systems that can interpret and act across different visual puzzles.

The practical implication is clear: recognising a bus is a weakening proxy for being human. Online trust increasingly depends on the wider interaction, account protections and layered defences. Better visual intelligence has changed both what assistants can accomplish and what websites can safely assume.

Assistants start doing the work

Another thread runs from our 2024 coverage of computer agents through Operator in January 2025 to this year’s discussion of Anthropic’s impact on IT services. The ambition was consistent: give software an objective and let it work across applications.

That ambition is now visible in products. OpenAI’s Astra release describes assistants researching sources, operating professional software, producing documents and testing websites. Anthropic’s Fable 5.1 release likewise centres extended coding and knowledge work. These are vendor descriptions of growing capability, not guarantees that any chosen workflow will succeed unattended.

Our reading is that the unit of value is changing from an impressive answer to a completed, checked task. Businesses will need to count the time spent reviewing and repairing work alongside the time saved creating it. Permissions, reliable records and recovery from mistakes become part of the product.

For workers, this also makes simplistic replacement forecasts inadequate. Automating part of a job changes the surrounding responsibilities. Who benefits depends on how the work is redesigned, who receives training and who remains answerable for the result.

Science makes progress at different speeds

The archive’s healthcare coverage stretches from detection and drug research to brain MRI triage and aged care. Here, the strongest evidence increasingly comes from testing a defined role in a real workflow.

Results published in January from Sweden’s MASAI randomised screening trial, analysing more than 105,000 women, found higher sensitivity with AI-supported mammography: 80.5%, compared with 73.8% for standard double reading. Specificity was unchanged. The trial did not establish a statistically significant reduction in cancers diagnosed between screening rounds. Its positive findings concern a particular screening arrangement, not universal diagnostic superiority.

A separate July study of NeuroVFM offers another instructive result. In a silent prospective evaluation of 1,155 CT and MRI studies, it achieved 92.6% balanced triage accuracy, yet missed critical findings in 21 of 155 affected patients. Silent evaluation means the system was being assessed without directing care. Both the performance and the misses belong in the story.

The same care applies to our coverage of robots, synthetic skin, new materials, photonic processors and quantum computing. A useful laboratory component can be a genuine advance while considerable engineering remains between it and a dependable product. Those stories collectively show a wider scientific toolkit; they do not establish that every promised application has arrived.

Trust still requires evidence

Our reports on erroneous legal citations and Deloitte’s flawed government report exposed a cost of treating fluent text as verified work. Stronger models make that distinction more consequential as their output becomes easier to accept.

The Federal Court’s April 2026 generative AI practice note makes responsibility concrete. It expects the responsible person to confirm, among other things, that cited legal authorities exist and support the propositions advanced. Professional accountability survives the arrival of a more capable assistant.

Creative work has followed a parallel path. Our early music copyright coverage and later Udio settlement story track the pressure for workable agreements. UMG’s announcement describes a settlement and licensing arrangements, not a court ruling resolving every training dispute. Australia’s government has ruled out a broad text-and-data-mining exception, keeping creator rights central to the discussion.

Our 2024 “nutrition labels” article anticipated another part of the response: content provenance. C2PA’s explanation draws an essential boundary. A record of an image’s origin and edits can help establish its history; it cannot, by itself, establish that the event depicted is true. Authenticity needs a more careful vocabulary as synthetic media improves.

The cloud acquired an address

Some of the clearest changes since our earlier coverage concern who controls access. June’s Fable 5 article described the abrupt suspension following US export controls. That story has moved on: Anthropic says the controls were lifted on 30 June and global Fable 5 access resumed on 1 July. Its September releases retain a distinction between generally available Fable and restricted Mythos access.

Restoration does not erase the lesson. Dependable access, contractual terms and jurisdiction can matter as much as model performance. For Australian organisations, sovereignty is a practical question about control over systems and data, including what happens when an overseas provider changes the conditions.

Energy deserves the same specificity. Our December environmental article emphasised efficiency and potential benefits. The IEA’s April 2026 update adds essential scale: worldwide data-centre electricity consumption reached an estimated 485 TWh in 2025, with around 950 TWh projected for 2030. Those totals cover all data centres, not AI alone. More efficient individual tasks can coexist with rising total demand and local grid pressure.

Canberra’s response is also developing beyond our May policy overview. The Office of AI was established on 15 July. On 26 August, National Cabinet agreed to work together on national AI laws, with Commonwealth legislation intended for early 2027, including protections around energy, water and data-centre development. These are forthcoming measures; announcing a rulebook and enacting it remain separate steps.

What the archive tells us now

The most persuasive change across these stories is the growing range of work AI can attempt. Reasoning, perception and software control are increasingly available together. Their combination opens possibilities that were harder to deliver when each capability stood alone.

Yet the archive also records what progress cannot settle automatically: who owns creative work, whose data is exposed, whether a clinical result transfers, who pays for infrastructure and how the benefits are shared. Those questions become more immediate as capability improves.

For FineSkyAi’s readers, the next milestone worth watching is dependable usefulness: a system that performs well on the actual task, leaves its work open to scrutiny and justifies its place in the workflow. There is plenty here to be excited about. The evidence is becoming more interesting than the predictions.

WhatsApp
LinkedIn
Facebook