A client's board asked for a slide on "our AGI strategy." The team building the slide had, at that point, a chatbot that summarised support tickets and a forecasting model that was still wrong about a fifth of the time. Nobody in the room could say what artificial general intelligence would actually need to do differently from those two systems for the word "general" to apply. The slide got made anyway. It said a lot of things. None of them were falsifiable.
We build data and AI systems at Shipshape Data, and this is one of the stranger ways a project stalls: not on the data, not on the model, but on a term that has quietly become a stand-in for "AI, but more impressive than what we currently have." Teams delay real decisions waiting for AGI to arrive, as though it were a product release with a date attached. Other teams brand a perfectly good narrow system as a step toward AGI because it sounds better in a deck. Both moves skip the actual question, which is what the system in front of you can and cannot do.
This guide covers what artificial general intelligence would mean if it existed, how it differs from the narrow systems running in production right now, why researchers who study this for a living cannot agree on a definition or a date, and what any of it should change about a business's AI plans this year. It will not tell you when AGI is coming, because nobody honestly knows, and anyone who states a precise year is telling you more about their incentives than about the technology.
Why the term gets thrown around loosely
"AGI" gets used in at least three different ways in the same conversation, often by the same person within the same sentence. Sometimes it means a system that matches human performance across most cognitive tasks. Sometimes it means a system that can learn a genuinely new task without being retrained for it. Sometimes it just means "very capable," a synonym for impressive that has drifted loose from any technical anchor.
A word doing three jobs
The looseness is not an accident. Vagueness is useful to a lot of people at once. It lets a lab describe its roadmap in terms grand enough to justify the funding round, without committing to a testable claim. It lets a vendor imply their product is closer to something remarkable than a careful reading of its capabilities would support. It lets a commentator warn about risk without specifying which risk, at which point, under which conditions. None of that is necessarily dishonest. It is just imprecise, and imprecision is comfortable because nobody can be proven wrong about a claim that was never pinned down.
For a business trying to plan, the imprecision is the actual problem. You cannot budget for, staff against, or govern a term three people in the same meeting mean differently. We have sat in planning sessions where "getting ready for AGI" was offered as a reason to delay a perfectly sensible narrow project, and nobody could say what readiness would actually consist of, or how anyone would know it had been reached.
What "general" is supposed to mean
Strip away the marketing and the core idea is reasonably simple. A general intelligence is one that can apply what it has learned to problems it was not specifically built or trained for, the way a competent adult can move from cooking to fixing a bicycle to reading a legal contract, drawing on the same underlying reasoning each time rather than needing a fresh brain for each domain. The bar most researchers actually mean, when pressed, is something like matching or exceeding human performance across the broad range of cognitive tasks a person can do, with the ability to transfer learning between them.
Transfer is the hard part
That transfer property is doing almost all the work in the definition, and it is exactly what today's systems struggle with. A model trained to be excellent at one thing rarely becomes good at an unrelated thing simply because it got better at the first. Human intelligence does this constantly and mostly without noticing: someone who learns to read a balance sheet does not start from zero when they later learn to read a tenancy agreement. That is the property people are actually pointing at when they say "general," even when they reach for vaguer language to describe it.
The AI running in your business today is narrow, and that's fine
Almost everything we deploy for clients is narrow AI: a model built and tuned for a specific task, trained on data relevant to that task, and evaluated against a specific metric for that task. A demand forecasting model is very good at demand forecasting and has no opinion on anything else. A document classifier sorts documents and does not, in any meaningful sense, understand the documents it sorts. Even a large language model answering open-ended questions is narrow in the sense that matters here: it is doing next-token prediction shaped by training, not reasoning from a general model of the world the way a person does, however fluent the output sounds.
Narrow does not mean weak
Calling something narrow can sound like a put-down, and it should not. A narrow system can still be extremely capable within its lane, sometimes more capable than a person, and that is usually all a business actually needs. Nobody wants their fraud detection model to also compose poetry. The narrowness is the design, not the limitation. Our longer look at narrow AI goes into how these systems are built and where they earn their keep, which is most places you will actually deploy AI in the next few years.
Generative models blur the picture a little
Large language models complicate the narrow-versus-general story because they operate across a genuinely wide range of tasks: drafting, summarising, coding, translating, holding a conversation about almost anything. That breadth is real and it is new. It is not the same thing as general intelligence, though, because breadth of surface task is not the same as the underlying transfer and reasoning that the AGI definition is reaching for. A model can write a competent legal summary and still fail at a change to the same problem that a ten-year-old would handle without pausing, because it has learned patterns in language rather than built a model of the situation the language describes. Breadth without robustness is still narrow, just narrow across more surfaces at once.
Why nobody agrees on the definition
Ask five researchers to define AGI and you will likely get five different tests, because they disagree about which capability actually matters most. Some anchor to economic impact: a system that can do the majority of economically valuable human labour. Some anchor to cognitive breadth: performance matching humans across a standard battery of tasks. Some anchor to something closer to autonomy: the ability to set and pursue its own goals over time without a human specifying every step. These are not the same finish line, and a system could conceivably cross one without coming close to another.
The moving-goalposts problem is real, and it cuts both ways
Critics of the field like to say the goalposts keep moving: chess was supposed to require general intelligence until a computer won, then it was Go, then it was passing a bar exam, and each time the milestone fell the definition quietly shifted to something further out. That criticism has some truth in it. But the shifting is not always bad faith. Some of those milestones turned out to be narrower than they looked when we first set them, achievable through pattern-matching and search rather than the general reasoning we thought they required. Moving the definition once you have learned that a supposed benchmark of general intelligence did not actually require general intelligence is not cheating. It is closer to correcting a hypothesis after evidence.
The argument about AGI is rarely an argument about the technology. It is an argument about which single number should count as proof of something nobody can fully specify in advance.
Definitions matter because money follows them
This is not an academic squabble. Contracts, funding rounds and even a famous corporate partnership agreement have reportedly hinged on a defined threshold for AGI, because whoever gets to say the threshold has been crossed can trigger real financial and legal consequences. When the definition affects who owes whom what, "we will know it when we see it" stops being an acceptable answer, which is exactly why the definition keeps getting fought over rather than settled.
Why timelines range from years to never
Survey a room of AI researchers on when AGI might arrive and you will get answers spanning decades, sometimes centuries, sometimes "not with current methods at all." That spread is not a sign that nobody has thought about it. It is a sign that the estimate depends heavily on assumptions nobody can currently verify: whether scaling current architectures keeps producing new capability at the same rate, whether some missing ingredient (a different kind of memory, a different kind of learning, something we have not identified yet) turns out to be required, and whether the definition being used is closer to the narrow economic-impact version or the far more demanding general-reasoning version.
Optimism has a business model
It is worth noticing who is making the confident short-term predictions. Organisations racing to build the most capable systems have an obvious incentive to say the milestone is close: it attracts investment, talent and attention. That does not make their predictions wrong. It does mean the source of a timeline is relevant information when you are deciding how much weight to put on it, in the same way you would weigh a growth forecast differently coming from the company selling the product versus an analyst with no stake in the outcome.
Pessimism has a business model too
The reverse also holds. Warnings of imminent, transformative risk generate attention and funding of their own kind, for research institutes, for advocacy, for regulation proposals. Neither side's incentives make their argument false. They do mean a stated timeline is rarely a neutral fact, and treating any single one as settled is a mistake regardless of which direction it points.
Progress on benchmarks isn't the same as progress toward AGI
Every few months a new model posts an impressive score on a reasoning benchmark, a coding test, or a professional exam, and the headline writes itself: a step closer to AGI. Some of that progress is genuine and useful. Some of it is the benchmark being gamed, deliberately or otherwise, because the training data increasingly overlaps with material the test was built from, or because the test rewards a pattern the model has learned to exploit without the underlying capability the test was designed to measure.
What a good score does and does not tell you
A model that scores well on a benchmark of logical puzzles has demonstrated it is good at that benchmark. Whether that generalises to the messy, underspecified problems a business actually has is a separate question, and it is the question that matters for planning. We have watched teams commission an expensive proof of concept on the strength of a headline benchmark score, only to find the model falls over on the specific, slightly unusual version of the task their business actually needs done. The benchmark measured something real. It was not the thing they needed measured.
What would actually change if it arrived
Suppose, for argument's sake, that a system meeting a serious definition of AGI existed tomorrow. What would actually be different for a business, as opposed to what gets claimed would be different? The honest answer is that a great deal of the work we currently do, cleaning data, defining what a customer record means, deciding which system owns which fact, would still need to be done. A more general reasoner does not know your business's definitions any better than a narrow one does. It still needs to be pointed at governed, trustworthy data to produce a trustworthy answer, and a smarter model given a mess produces a more articulate mess, not a correct one.
What might genuinely shift
What could plausibly change is the amount of task-specific engineering required to get a system working on a new problem. Today, moving a model from one use case to a meaningfully different one usually takes real integration work: new prompts, new evaluation, sometimes new fine-tuning. A system with genuine transfer ability would need less of that per new task, and a smaller integration bill is a genuine commercial benefit worth planning around once it is real rather than promised. That is a useful shift if it happens. It is not the same as the system suddenly needing no oversight, no governance and no checking, and treating it that way would be a mistake regardless of how capable the underlying model became.
Autonomy is a separate question from capability
A more general system is also not automatically a more autonomous one. Capability and the amount of independent authority a system is given are different dials, set by the people deploying it, not by the model itself. The risk management discipline that applies to a narrow model today, knowing what it is allowed to decide unsupervised and what it must escalate, does not become optional just because the underlying model got more capable. If anything, a more capable system with the same weak oversight is a bigger version of the same problem, not a different one.
Where to start
Do not build a plan around a date nobody can honestly give you. Build one around the capability gaps you can actually name: which tasks in your business are too varied for your current narrow tools, which ones would benefit from a system that transfers better between related problems, and which ones are perfectly well served by the narrow model you already have, however unglamorous that sounds in a board pack. Write those gaps down in plain language, with an example of the specific task that currently fails, and revisit the list each time a new model release makes big claims. Most releases will not close the gaps you actually wrote down, and knowing that in advance saves a lot of wasted evaluation time.
Keep governance proportionate to what a system is actually allowed to do, not to what it is rumoured to be capable of in some future version. A useful discipline here is treating every new model the same way you would treat a new hire with an unverified reference: give it a narrow, checked scope first, and widen it only as it earns that trust with evidence, not enthusiasm. That habit holds regardless of whether the underlying models get more general in the next five years or stay roughly where they are.
Most of the actual return on AI in the next few years will come from applying narrow systems well against a data foundation that can support them, the unglamorous work that enterprise AI programmes tend to underestimate at the planning stage and then rediscover the hard way, mid-project. If your business is trying to work out what its AI plans should actually depend on, rather than what a slide about AGI implies they should depend on, talk to us. We would rather help you scope this year's real capability gap than help you write a strategy slide for a date nobody can give you.