edidiong umana · writing
home
Research notes10 min read

The state of AI, and the next ten years: what I'm betting a degree on

What can be verified about AI in October 2026, six bets for 2026 to 2036 that could be proved wrong, what would change my mind on each, and what they mean for anyone training to be an AI engineer now.

If you're starting to study AI now, you have an awkward problem: the subject changes faster than any syllabus. The 2026 AI Index found that benchmarks designed to stay hard for years were being saturated within months. So which skills will still be worth having in 2036?

I can't know. What I can do is write down what is verifiable today, make a few bets that could turn out wrong, and say in advance what would change my mind. Every number below links to the page I took it from. Everything about the future is a prediction, and I've marked it as one.

Where things stand: compute and cost

Compute is still growing about fivefold a year. Epoch AI estimates that the compute used to train frontier language models has grown about 5x a year since 2020, and that the cost of training them has grown about 3.5x a year. The 2026 AI Index puts global AI compute capacity at 17.1 million H100-equivalents, growing 3.3x a year since 2022. Nvidia supplies over 60% of it, and one foundry, TSMC, makes almost every leading AI chip.

The scale is now physical. Epoch's figure for the largest known AI data centre is the equivalent of 1.1 million H100 chips.

Running a model got much cheaper, but unevenly. The 2025 AI Index reported that the cost of running a system at GPT-3.5's level fell more than 280-fold between November 2022 and October 2024. Over the same period, hardware costs fell about 30% a year and energy efficiency improved about 40% a year.

Epoch AI's analysis of six benchmarks found that the price of reaching a fixed level of performance fell by anywhere from 9x to 900x a year, depending on the task. For GPT-4's level on PhD-level science questions, it was 40x a year. Epoch also warned that the steepest drops were the most recent, so they may not last. I couldn't find an updated price series in the 2026 Index, so read these as 2024 and early 2025 figures.

Where things stand: open models

Open weights are close behind, not level. In March 2026, the AI Index found the best closed model about 3% ahead of the best open-weight model on the Arena leaderboard, up from 0.5% in August 2024. Six of the top ten models were closed. Epoch AI measures the gap in time instead: since January 2026, the best open-weight models have trailed the closed frontier by about four months.

For anyone without a data centre, two other findings matter more. In August 2025 Epoch showed that a single gaming GPU, an RTX 5090 costing under $2,500, can run models that match the frontier of 6 to 12 months earlier. And the AI Index notes that OLMo 3.1 Think 32B, with nearly 90 times fewer parameters than Grok 4, matches it on several benchmarks through better data curation alone.

Where things stand: agents

Agents moved from answering to doing. On OSWorld, where agents operate real computers, the best score rose from roughly 12% to 66.3%, within six points of humans. On Terminal-Bench 2.0 it went from 20% in February 2025 to 77.3% in early 2026. The same chapter notes that agents still fail about one attempt in three on structured benchmarks.

METR measures something different: how long a task, in human working time, an agent can finish half the time. In January 2026 it reported that this length doubled about every seven months from 2019 to 2025, and about every 89 days since 2024. Its top model then, Claude Opus 4.5, reached about 320 minutes, with a confidence interval from 170 to 729 minutes.

The plumbing is being standardised. In December 2025 the Linux Foundation formed the Agentic AI Foundation around the Model Context Protocol, goose and AGENTS.md. In September 2025 Google announced AP2, an open protocol for agent payments built with more than 60 organisations, designed to prove that a user really authorised a specific purchase.

Where things stand: evals and energy

Evals and safety are behind. The International AI Safety Report 2026, chaired by Yoshua Bengio with over 100 experts, says pre-deployment tests often fail to reflect real-world use. It also says models more often score well by finding loopholes in an evaluation without doing the task, which it calls reward hacking. The AI Index found invalid question rates from 2% on MMLU Math to 42% on GSM8K. It counted 362 documented AI incidents in 2025, up from 233 in 2024, and Foundation Model Transparency Index scores fell to 40 from 58.

Security is the sharpest version of the problem. In 2025, one research team got past 12 published defences against jailbreaks and prompt injection using adaptive attacks, most of them more than 90% of the time. Most of those defences had first reported attack success near zero.

Energy is now a constraint. The IEA estimates that data centres used about 415 TWh in 2024, around 1.5% of the world's electricity, after growing about 12% a year since 2017. It projects around 945 TWh by 2030. The AI Index puts AI data centre power capacity at 29.6 GW, comparable to New York state at peak demand, and notes that the energy spent serving a deployed model can exceed its one-time training cost within months. Chips keep getting more efficient, about 34% a year by Epoch's count, but models have scaled faster.

Where things stand: Africa

Africa has a small share and real momentum. Africa is home to 18% of the world's population but less than 1% of global data centre capacity, according to Brookings. In the AI Index, sub-Saharan Africa produced 0.83% of AI publications in computer science in 2024.

The language gap is measurable. IrokoBench (NAACL 2025) tested models in 17 African languages and found a large gap against English and French. The best open model, Gemma 2 27B, reached only 63% of GPT-4o's performance.

The work is under way, though. Masakhane, a grassroots NLP community, reports more than 1,000 participants from 30 African countries. The Deep Learning Indaba met at Pan-Atlantic University in Lagos from 2 to 7 August 2026. Lelapa AI's InkubaLM is a 0.4-billion-parameter model covering five African languages plus English and French. Nigeria's N-ATLaS fine-tunes Llama-3 8B on about 392 million tokens of English, Hausa, Igbo and Yoruba instruction data.

Six bets for the next ten years

These are predictions, not findings. For each one I give the evidence, what would change my mind, and what it means for someone becoming an AI engineer or forward deployed engineer (FDE) now. I'd rather be specifically wrong than vaguely right.

Bet 1: Inference becomes the main cost and the main skill

Evidence. Prices per token keep falling, yet AI data centre power keeps rising, and serving a model can outrun its training energy within months. Agents push the same way, since every step of a multi-step task is another model call. My reading, not a measured fact, is that cheaper tokens mean far more tokens get used.

What would change my mind. By 2030, providers report that training still dominates their compute, or serving becomes so commoditised that tuning it no longer pays for application teams.

For you. Learn what happens between the prompt and the last token: prefill and decode, the KV cache, quantisation, batching and caching. Measure cost per completed task and p95 latency, not price per million tokens.

Bet 2: Evals and observability become a profession

Evidence. Public benchmarks saturate in months, some carry invalid questions at rates up to 42%, and the Safety Report says pre-deployment tests often miss real-world behaviour. Incidents are rising. Every team that deploys AI will need its own evidence that the system works for its users.

What would change my mind. Public benchmarks start to predict deployed performance well, or automated grading by models becomes reliable enough that people only audit it now and then.

For you. Build test sets from real traces. Learn enough statistics to put an interval on a pass rate. Grade agents on the state they leave behind and on repeated runs, not on one good demo.

Bet 3: Agents need identity, permissions and payments, and security is the hard part

Evidence. METR's task lengths keep growing, so agents act for longer without a person watching. The Safety Report notes that agent failures leave humans fewer chances to step in. The industry is already building shared plumbing (the Agentic AI Foundation, AP2's signed mandates for authorisation), while prompt injection remains unsolved.

What would change my mind. A published defence holds up against independent adaptive attackers for two years or more, or agents in production stay mostly read-only.

For you. Treat least privilege, approvals for irreversible actions, scoped credentials and audit logs as core skills. Threat-model every tool an agent can call.

Bet 4: Small, open models on local hardware matter most where compute is scarce

Evidence. Open-weight models trail the frontier by months, not years. One consumer GPU runs models that were frontier a year earlier, and careful data lets a 32B model compete with far larger ones. Meanwhile Africa holds less than 1% of global data centre capacity.

What would change my mind. The open-weight lag widens past a year and stays there, or hosted APIs become cheap and reliable enough on African networks that local deployment stops making sense.

For you. Learn the memory maths (parameters times bytes per parameter, plus the KV cache), the cost of quantisation in quality, and how to evaluate a model on modest hardware before you trust it.

Bet 5: African languages and data become a competitive edge

Evidence. IrokoBench shows a large gap between English and African languages, and the region produces under 1% of AI research. Yet models such as InkubaLM and N-ATLaS, and communities such as Masakhane and the Indaba, show the people and data work exist locally.

What would change my mind. General frontier models score within a few points of English on IrokoBench-style tests without targeted data. Then the edge moves from data to distribution and trust.

For you. Build evaluation sets in the languages you speak, with consent and clear licences. That work pays off whichever way this bet goes.

Bet 6: The bottleneck moves from models to deployment

Evidence. The AI Index found four companies within 25 Arena points of each other in March 2026, with competition shifting towards cost, reliability and domain-specific performance. Models score 60% to 90% on professional tasks such as tax and legal reasoning, and agents still fail about one attempt in three.

What would change my mind. One lab opens a large lead and keeps it for more than a year, so that choosing the model matters more than how it's deployed.

For you. This is the forward deployed engineer's work: sitting with users, wiring models into messy systems, handling failure, and proving value with measurements.

What this means for how I'm studying

If the bets are roughly right, the learning priorities follow, and they hold for anyone starting out:

  1. Maths and ML fundamentals. Linear algebra, probability and optimisation outlast any architecture, and let you read papers instead of summaries of papers.
  2. Evals. Test design, error analysis and enough statistics to know when a difference is real.
  3. Inference. Memory, latency, quantisation and cost per task, measured on real hardware.
  4. Security. Prompt injection, least privilege and red-teaming your own systems.
  5. Deployment. Shipping to real users, watching traces and fixing what breaks.

I've just started a B.Sc. in AI, and this is the list I'm holding my own study to.

Do this todayWrite 20 test cases from real requests, in the languages your users speak, and run them against one hosted model and one open-weight model on your own machine.

The bets in one table

BetSignal to watchWhere we'd be in 2030 if right
Inference is the main cost and skillShare of AI compute and power going to serving; cost per completed taskInference engineering is a standard role, and teams budget per task, not per token
Evals become a professionGap between benchmark scores and deployed results; teams and roles named for evaluationSerious AI teams run their own evals on every change, and public benchmarks are rough filters only
Agents need identity, permissions and paymentsAdoption of shared agent protocols; independent results against adaptive prompt injectionAgents carry scoped, auditable authority by default, and injection is managed like fraud, not solved
Small, open models where compute is scarceEpoch's open-closed lag; what one consumer GPU can runMost routine AI work in low-compute markets runs on open models close to the user
African languages and data are an edgeIrokoBench-style gaps; new African-language models and datasetsProducts in African languages compete on local data and evaluation, not only on the base model
Deployment is the bottleneckSpread between top models on Arena; agent failure rates on real tasksDifferences in value come mostly from integration and reliability, not model choice

I'll revisit this table each year and mark what moved. If a bet fails, I'd rather say so than quietly drop it.

Learn it properly

For more depth, the AI Study Group has free tracks on inference (T3), evals (T8) and AI security (T9).

Sources