// listen — narrated by openai tts-1-hd, voice: echo
// primary blocked? play from mirror ↗
Most of the people who would benefit most from a working AI assistant are not in the room where the assistant is built. This is not an accusation. It is an observation about how the field is structured, and the consequences of that structure are not visible from inside it.
Modern frontier AI is a real resource. It is also delivered to the world through a specific architecture: a small number of laboratories in a small number of cities, serving everyone through metered APIs over reliable bandwidth, billed in dollars at the standard rate. That architecture has unstated assumptions baked into it. Each assumption screens out a different population. Stacked, they leave the default product structurally mismatched to much of the world it could otherwise reach.
I write from a chip company in San Diego, and I grew up in a small town in Odisha. I know the people the architecture works least well for, because they are the people I grew up around. They are not case studies. The field has interesting unfinished engineering work to do here, and the people best positioned to know what AI should do for a given population are people from that population.
This post is not a call for regulation, foreign aid, or government intervention. It is a description of an architecture, and of the population it leaves out, in the hope that someone reads it and decides to build something.
The numbers
A single LLM query is cheap to serve, and getting cheaper. A 2025 measurement study covering more than 32,500 measurements across 21 GPU configurations and 155 model architectures gives a systematic basis for inference energy estimation [1]. Independent testing on an 8×H100 node, running Llama 3.3 70B in FP8 with vLLM at batch size 128, reported roughly 0.39 joules per token (input plus output combined) [2]. Order-of-magnitude estimators put a typical short GPT-4o-class exchange at about 0.3 to 0.43 Wh, with the caveat that these are working numbers from public-data extrapolation rather than vendor disclosure [2, 3]. Reasoning models cost considerably more per query, because they emit thousands of internal tokens before producing the visible answer; published estimates put the multiplier at roughly 30 to 50 times, depending on workload [3]. Training a frontier model is a separate matter, and runs into the high tens to hundreds of millions of dollars per state-of-the-art run [4]. Training is amortised across all users. The cost that gets passed through is inference, and inference is no longer the limiting factor.
The retail price of a consumer subscription has converged. ChatGPT Plus, Claude Pro, Google AI Pro, and Perplexity Pro all sit at twenty dollars a month [5, 6, 7]. Premium tiers cluster between one hundred and two hundred and fifty dollars: Claude Max, ChatGPT Pro, Google AI Ultra. OpenAI launched ChatGPT Go in India on 19 August 2025 at ₹399 a month, expanded it through Asia and Europe over the following months, and made it globally available on 15 January 2026 at $8 a month, with billing on local rails such as UPI, GoPay, and Pix [8].
Now place that against income.
| Country | Median monthly individual income | ChatGPT Plus list price | Plus as % of median income | Source |
|---|---|---|---|---|
| United States | ~$5,000 | $20 | ~0.4% | BLS Q4 2024 |
| Germany | ~€24 incl. VAT | ~0.7% | UNECE / Destatis | |
| Brazil | R368) | ~R$100 | ~5.3% | IBGE PNAD 2024 |
| India | ₹12,500 (~$134) | ₹1,999 (~$13)* | ~16% face / ~50%+ at PPP | PLFS 2024 |
| Indonesia | $20 | ~10% | BPS 2024 | |
| Nigeria (Lagos median) | 20‡ | 18-57% | NBS / Forbes |
* India price via the Apple iOS track. The web list price defaults closer to 6.40 due to regional pricing; the web default is closer to $20.
Methodology note. Income figures are nominal monthly individual earnings in the most recent year for which the cited source publishes them; definitions and survey methodologies differ. The Indonesia row is an average rather than a median, and the Nigeria row is a Lagos metro estimate rather than a national figure. Subscription prices are list face prices in local currency at the time of writing; some plans (Apple iOS billing, regional rollouts, promotional bundles) resolve lower in practice. PPP-adjusted percentages would, in most rows here, widen the gap rather than close it.
The Lagos number is the one that lingers. The median Lagos worker, by a Forbes-reported survey of Lagos-region employees, earns less per month than the federal minimum wage of ₦70,000 [9, 10]. On the standard web price, ChatGPT Plus is more than half of that. The Indian face-price gap looks smaller until you adjust for purchasing power, at which point a thirteen-dollar subscription corresponds to roughly seventy to ninety US dollars in equivalent local purchasing power, depending on whether the GDP or household-consumption PPP factor is used.
The architecture has noticed. ChatGPT Go at eight dollars lowers the floor and runs on local rails. Google’s AI Plus launched in India at ₹399 a month with an introductory price of ₹199 for the first six months [11]. These are real moves, and bounded ones. A thirteen-dollar face price is still around ten percent of Indian median monthly income, and the strongest reasoning models, the longest contexts, and the heaviest tool use remain on the hundred- and two-hundred-dollar plans.
There is one more piece of data that belongs here, which is the unit economics. Heavy users on flat-rate plans cost more than they pay. Anthropic introduced weekly rate limits on its two-hundred-dollar Max-20 tier in 2025 [12]. The standard twenty-dollar plan exists because the platform expects most subscribers to barely use it. If the median user in a tier-2 Indian town actually used it the way a Bay Area developer does, the math at the current price would not hold.
A reasonable reply at this point: consumption is distributing fast, and faster than this section suggests. The World Bank reports that more than forty percent of global ChatGPT traffic in mid-2025 originated in middle-income countries, led by Brazil, India, Indonesia, and Vietnam [13]. Microsoft’s AI Economy Institute estimates global generative-AI adoption at sixteen percent of world population by late 2025 [14]. The technology is reaching those users. The argument of this post is not that it is not. The argument is that what it reaches them as, at the price they can pay, in the languages they speak, on the networks they have, with the trustworthiness their work demands, on infrastructure their grids can carry, is a categorically different product from the one a Bay Area engineer is using to write code. Consumption distributes. Definitional power does not.
Price, in any case, is the easiest of the exclusions to fix. The harder ones are technical.
Five exclusions
Suppose you fix price. Make it free, no rate limit, the highest tier for everyone. The next problem is capability, and the gap is not the one most readers expect. Free-tier base models are reasonably strong in 2026: Anthropic serves Sonnet 4.6 to free users, OpenAI’s free tier runs on GPT-5.3 [5, 6]. The free-versus-paid gap is now mostly about reasoning depth, usage caps, tool access, and longest-context handling. The gap that matters more for the populations this post is about is multilingual. On MMLU-ProX, the best-performing model scored 70.3 percent on English and 40.1 percent on Swahili, a thirty-point gap on a single benchmark for a single model [15]. AfroBench, covering sixty-four African languages across fifteen tasks, characterises the gap between English and African languages as “large” across most tasks even on GPT-4o and Gemini 1.5 Pro [16]. On IrokoBench, the best open-weight model scored only fifty-eight to sixty-three percent of the closed-model score on African-language tasks [17]. On IndicGenBench, across twenty-nine Indic languages, PaLM-2-L scored 83.7 in English and 41.1 on certain Indic generation tasks [18]. The gap is twenty to forty points on the things that matter most for low-resource speakers.
Suppose capability is also fixed. Assume the model speaks Hindi as well as English. Bandwidth and latency are the next wall. Median mobile download speed is roughly 29 to 43 Mbps in Indonesia and 43 Mbps in Nigeria, both numbers concealing very wide rural-urban gaps inside each country [19, 20]. Rural broadband penetration in India is 29.3 percent against 93 percent urban [21]. The ITU’s mobile-broadband basket cost 2.74 percent of GNI per capita in Nigeria in 2025 against 0.68 percent in the US, an order-of-magnitude affordability difference for being connected at all [22]. And LLM streaming over weak networks degrades disproportionately. An academic study showed that under unstable network conditions, ChatGPT’s TCP-based token streaming exhibited stall ratios that a custom protocol could reduce by 71 percent, because a single lost packet blocks all subsequent token rendering even when newer-token packets have already arrived [23]. Common industry estimates put each additional second of perceived chatbot delay at roughly seven to ten percent more user abandonment, though the figure varies by study and is best treated as an order-of-magnitude rule [24]. The product over low connectivity in rural Bihar is not a slower version of the product over fibre in San Francisco. It is a different product.
Suppose bandwidth is also fixed. Trust is the next wall, and it is the most dangerous of the four, because it fails silently. A small-town lawyer in Lucknow drafting bail petitions in Hindi will get a confident answer about Indian inheritance law that sounds correct and is wrong. Hemrajani (2025), evaluating GPT-4 and Claude on Indian legal practice from the National Law School of India, found the models match or exceed junior lawyers at drafting and issue-spotting but produce frequent fabrications on specialised Indian legal research [25]. Stanford RegLab’s hallucination work puts hallucination rates at sixty-nine to eighty-eight percent on specific US legal queries, with subsequent work finding the rates worsen for queries about Sub-Saharan African or Indic legal contexts [26, 27]. The RegLab paper’s own warning is that legal hallucinations may “have the opposite effect on the justice gap by systematically disadvantaging those who need the most help” [26]. The lawyer in Lucknow is exactly the user the warning is about. So is the community health worker in rural Sulawesi answering questions about post-natal care, and the municipal clerk in Lagos digitising decades of land records. None of them has a way to know that a confidently stated answer is wrong. Rest of World tested GPT-3.5 in Tigrinya and Amharic in 2023 and got back gibberish; their February 2026 retrospective confirms Bengali, Swahili, Urdu, and Thai are still where the model breaks [28, 29].
Suppose all four are fixed. Energy is the last wall, and it is where the architecture itself fails, not just the user-side experience. A hyperscale data centre campus draws hundreds of megawatts. The US hosts roughly 5,427 data centres, more than ten times any other country, and accounts for about half of all hyperscale capacity worldwide [30, 31]. McKinsey’s 2030 outlook projects 156 to 219 GW of global data-centre capacity, against roughly 47 GW today, with about seventy percent of new demand AI-driven (a projection, not a measurement) [32]. In Virginia, data centres already consume around twenty-six percent of state electricity. In Ireland, the Central Statistics Office reported that data centres accounted for 22 percent of all metered electricity in 2024, up from 5 percent in 2015, with EirGrid forecasting roughly 30 percent by 2032 [33, 34]. Hyperscalers go where there is cheap reliable power, abundant water for cooling, and stable politics: Northern Virginia, Iowa, the Nordics, now Riyadh. They do not go to Lagos, where the national grid collapsed nine times in 2024, where the average household receives roughly four hours of power a day, and where ninety-six percent of industrial electricity comes from private generators [35, 36, 37]. The Lagos clerk is not waiting for ChatGPT to be cheaper. He is waiting for the lights. In one Prayas Energy Group monitoring window, rural Uttar Pradesh averaged nine hours of daily outages while the official figure put the shortfall under one percent: the gap between what the grid is reported to be doing and what it actually does is itself part of the architecture [38]. As long as inference happens on a Virginia campus, the user in Sulawesi pays for the geography in latency, in reliability, in trust they should have but cannot.
By the time you reach the last wall, no patch on the existing architecture solves the underlying problem.
The view from inside
I started a PhD in 2018, before GPT-2. The lab I joined did energy-harvesting sensors, intermittent computing, and computational storage: the systems engineering of doing useful machine work at the edge of the network, on devices that may or may not be powered, on memories that lose state when the battery runs out. The questions we worked on were small and pragmatic. How do you train a neural network on a microcontroller that goes dark every twelve seconds? How do you do continuous learning on a solar-powered server in a field, with intermittent connectivity, without losing weeks of data on a power outage? The papers I am most proud of (NExUME on intermittency-aware DNN training, Salient Store on computational storage for continuous learning, Usás on battery-free continuous learning on solar-powered edge servers) all live inside this set of constraints.
Over the seven years of the PhD, the field’s centre of gravity moved. The interesting question stopped being “how do we make ML fit on tiny devices” and became “how do we spend a hundred million dollars on a training run.” This is not a complaint. The scaling bet was a real bet, and it paid off; the resulting models do things small models do not. But the field also made a choice about where to put its attention. The systems work that would have mattered most for accessibility (energy harvesting, intermittent and federated learning, on-device inference, computational storage, the entire toolkit for doing useful work without a data centre behind you) got proportionally less attention as the rest of the field scaled. Not zero. The work continued and continues. But the talent and the funding and the prestige flowed elsewhere.
The consequence is that the delivery architecture of modern AI was built on an assumption of abundance: abundant electricity, abundant cooling water, abundant bandwidth, abundant compute, abundant English text, abundant capital. None of those assumptions hold in most of the places where this technology would matter most. Most of the world does not live in that architecture, and the architecture was not designed for it.
This is a real resource. The question is how it gets to more people, and the answer is not just cheaper subscriptions.
I do not think this is anyone’s fault. The frontier labs built something real, in the constraints they had. But the labs are not best positioned to fix the distribution. The work that needs to happen next is closer to the constraints than to the frontier, it is systems work, and most of it is uncashed.
What the room does not know
An engineer in San Francisco does not have ground truth on what AI should do for the lawyer in Lucknow drafting bail petitions in Hindi. She does not know what the failure modes look like for the health worker in rural Sulawesi answering questions about post-natal care from a phone with two bars of signal, or for the clerk in Lagos digitising decades of land records under a four-hour power schedule. This is not a moral failing of the engineer. It is a property of geography and life experience. The knowledge required to build for these users does not exist in the room where the model is trained, and it cannot be acquired by querying a focus group.
The implication is uncomfortable but not radical: the work of figuring out what AI should do for a given population is local work. It belongs to people who live in the places they are building for. AI4Bharat at IIT Madras is doing it for twenty-two Indian languages, with a dataset of twelve thousand hours of speech from over twenty thousand speakers across more than two hundred districts [39]. Masakhane is doing it across the African continent through a network of more than two thousand researchers [40]. Lelapa AI in Johannesburg has shipped a 0.4-billion-parameter model called InkubaLM for five major African languages, on the explicit reasoning that more than seventy percent of African smartphone users are on entry-level devices and heavy compute-intensive models will not run there [41]. SEA-LION at AI Singapore covers eleven Southeast Asian languages [42]. Karya in Bengaluru has built a data-collection model that pays workers around five dollars an hour and routes royalties on resale back to them [43]. Cohere For AI’s Aya covers a hundred and one languages [44]. These are not adjacent to the frontier-AI story. They are the part of the story the frontier cannot do.
Close
The argument of this post is only worth making if there is work to do. There is.
The next post asks what a solar-powered local AI appliance would actually look like in numbers: power budget, storage budget, model sizes that fit, the workloads it can serve, what the device can and cannot honestly do. After that, posts on small models for narrow tasks, CPU-friendly inference, low-cost retrieval pipelines, edge deployment for local languages, and energy-aware serving. Build logs of things that work, and things that do not. Paper reads through a systems lens, in the corner of the literature where the real work is happening, mostly without much attention.
I do not have a thesis for this blog beyond the one I have already given. This is a real resource. The architecture by which it reaches the world is mismatched to most of its users. The engineering work to fix that is interesting and largely undone. The work that needs to happen next is closer to the constraints than to the frontier.
Bibliography
- Castro, R. et al. (2025). From Prompts to Power: Measuring the Energy Footprint of LLM Inference. arXiv:2511.05597. Large-scale measurement study covering 32,500+ measurements across 21 GPU configurations and 155 model architectures. [Preprint.]
- Lin, L. H. (2025). Internal User Test: Llama3-70B Inference Efficiency on H100, gist bf81a9c7dfc4244c974335e1605dcf22; summarised at llm-tracker.info. ~0.39 J per token (input + output) on 8×H100 with FP8 and vLLM at batch 128. [Independent test.]
- Per-query energy consumption of LLMs (Feb 2026). muxup.com. [Independent estimator extrapolating from public data.]
- Epoch AI (June 2024). “How much does it cost to train frontier AI models?” epoch.ai/blog.
- Anthropic (Feb 2026). “Introducing Claude Sonnet 4.6”; Anthropic pricing page (accessed April 2026).
- OpenAI Pricing Page (accessed April 2026). chatgpt.com/pricing.
- FindSkill.ai (April 2026). “AI Pricing Compared 2026.” [Secondary aggregator; cross-checked against vendor pages where possible.]
- OpenAI (15 Jan 2026). “Introducing ChatGPT Go, now available worldwide” (openai.com/index/introducing-chatgpt-go); ChatGPT release notes (19 Aug 2025 entry, India launch at ₹399/month); OpenAI multi-currency billing help page (accessed April 2026).
- National Bureau of Statistics Nigeria (2024) on national minimum wage of ₦70,000 (effective July 2024); Forbes / Time Doctor reporting on a 2024 Lagos-region employee survey finding median monthly earnings of ~₦60,000, below the federal minimum.
- Turisvpn (Feb 2026). “Where to Find the Cheapest ChatGPT Plus Country.” [Secondary; used only for regional-pricing facts.]
- Times of India (Nov 2025). “Google launches AI Plus subscription in India at introductory price of Rs 199 per month.”
- Adewale, A. (March 2026). “The Anthropic Irony.” Medium, citing The Information (Jan 22, 2026). [Secondary; primary report behind paywall.]
- World Bank (Nov 2025). Digital Progress and Trends Report 2025.
- Microsoft AI Economy Institute (Jan 2026). Global AI Adoption in 2025.
- Xuan, W. et al. (2025). “MMLU-ProX: A Multilingual Benchmark for Advanced LLM Evaluation.” arXiv:2503.10497.
- Ojo, J. et al. (2025). “AfroBench: How Good are Large Language Models on African Languages?” arXiv:2311.07978v5 (ACL Findings 2025).
- Adelani, D. et al. (2024). “IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models.” arXiv:2406.03368; NAACL 2025 follow-up.
- Singh, S. et al. (2024). “IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages.” ACL 2024 (arXiv:2404.16816).
- Speedtest / Ookla via Statista (Oct 2024). Indonesia mobile-network speed.
- Speedtest / Ookla (Q4 2025). Mobile Internet speed by country.
- Shruthi K.A. et al. (2021/2024). “A Survey on Rural Internet Connectivity in India.” arXiv:2111.10219, citing TRAI.
- ITU DataHub (2025). “Mobile broadband data and voice low-consumption basket total.”
- Robust Transport for LLM Token Streaming under Unstable Network. arXiv:2401.12961 (2024).
- AIonX (2026). “AI chatbot response time benchmarks.” [Industry blog; figure used as order-of-magnitude rule only.]
- Hemrajani, R. (2025). “Evaluating the Role of Large Language Models in Legal Practice in India.” arXiv:2508.09713 (National Law School of India).
- Dahl, M. et al. (2024). “Hallucinating Law: Legal Mistakes with Large Language Models are Pervasive.” Stanford HAI/RegLab.
- Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries. arXiv:2511.06700 (2025).
- Rest of World (2023). Andrew Deck investigation on ChatGPT in Tigrinya and Amharic.
- Rest of World (Feb 2026 retrospective). Persistence of low-resource language gaps.
- Stanford HAI (2026). 2026 AI Index Report.
- Network Installers (2025). “AI Data Center Statistics & Trends,” summarising Synergy Research. [Secondary.]
- McKinsey & Company (2025). AI Data Center Capital Expenditures Outlook 2025-2030. [Projection, not measurement.]
- Pew Research Center (Oct 2025). “US data centers’ energy use amid the AI boom,” citing IEA 2024.
- Central Statistics Office Ireland (10 June 2025). Data Centres Metered Electricity Consumption 2024. cso.ie. EirGrid 2032 forecast cited via the same release.
- BusinessDay NG (Jan 2026). “Blackout as Nigeria’s power grid collapses for first time in 2026”; reports on 2024 collapses.
- Wikipedia, “Nigerian energy supply crisis,” summarising primary sources.
- Georgetown Journal of International Affairs (Aug 2024). “Illuminating Nigeria: Grid and Off-Grid Electricity.”
- The Leap Blog (Sept 2025). “A Review of Outage Reporting by Indian DISCOMs,” citing Prayas (Energy Group) Electricity Supply Monitoring Initiative (ESMI). The nine-hour rural / 0.9 percent official contrast comes from a January 2017 ESMI sample and is illustrative of the official-vs-monitored gap rather than a current state-wide daily average; ESMI tracks ~160 locations across 14 districts.
- AI4Bharat. ai4bharat.iitm.ac.in. IIT Madras / Microsoft Research India.
- Masakhane. masakhane.io; Carnegie Endowment (Dec 2025) on Masakhane network scale.
- Lelapa AI (Aug 2025). “Locked Out by Design”; Lelapa AI / Zindi Buzuzu-Mavi Challenge press release (June 2025) for “more than 70% of smartphone users on entry-level devices and internet penetration in Sub-Saharan Africa still hovering around 33%”; InkubaLM model card and arXiv release.
- Susanto, Y. et al. (2025). “SEA-LION: Southeast Asian Languages in One Network.” arXiv:2504.05747 (AI Singapore).
- Microsoft Asia (2024). “Village by village…”; Time (2023). “The Indian Startup Making AI Fairer.”
- Aryabumi, V. et al. (2024). “Aya 23: Open Weight Releases to Further Multilingual Progress.” arXiv:2405.15032 (Cohere For AI).