Narrow AI is an AI system designed for a defined kind of work. An email spam filter classifies messages; a program trained to play Go chooses moves. Either can be highly capable at its task. Neither result, by itself, shows that the system can handle unrelated work with the same reliability.
“Narrow” describes the scope of a system’s intended and demonstrated ability, not whether its technology is simple. It also does not mean the system follows only hand-written rules. A narrow system may learn patterns from data, make predictions on new inputs, and improve when its developer updates it.

What does “narrow” mean in AI?
The OECD’s definition of an AI system covers systems that infer from inputs to produce predictions, content, recommendations, or decisions. Narrow AI is a useful name for a system whose intended and tested uses are limited to a particular domain or set of tasks. “Weak AI” and “artificial narrow intelligence” are also used for this general idea. “Weak” does not mean ineffective.
The scope may be broad enough to handle many variations within a task. A spam filter can assess unfamiliar messages without being given a rule for every possible email. The boundary appears when someone asks it to do a different job or assumes its accuracy will carry over to new conditions. For an introduction to the wider field, see our plain-English AI guide.
Real examples of narrow AI
Spam detection. Google describes Gmail’s spam filtering as a combination of AI-driven filters and user feedback. The system considers signals such as sender information and authentication to classify incoming mail. Its job is to identify suspicious or unwanted messages. A legitimate message can still be marked as spam, and unwanted mail can still get through; users can correct a classification.
Game playing. Google DeepMind’s AlphaGo combined neural networks and search to play Go at an extraordinary level, defeating Lee Sedol in 2016. This is a historical example of exceptional performance in a defined domain. Winning at Go does not establish competence at interpreting a contract or planning a medical treatment.
A legal-workflow illustration. Suppose a law firm builds a classifier to flag indemnity clauses in contracts for a lawyer’s review. If trained and tested well, it could help staff find candidate passages in many documents. It would not, on that evidence alone, decide whether a clause is enforceable, identify every unusual allocation of risk, or replace review by counsel. This is a hypothetical example, not a claim about a particular commercial product.
Narrow AI compared with general AI
The difference is breadth, not a guarantee of quality. A narrow system is intended for specific work and should be evaluated on that work. Artificial general intelligence, or AGI, refers to a much broader goal: a system able to perform well across many kinds of intellectual tasks and transfer what it learns between them. Researchers use different definitions and proposed thresholds for that goal; success on one task does not settle the question. See our AGI guide for the competing definitions and examples.
Modern AI assistants can draft, summarize, analyze images, use tools, and switch among tasks. That breadth makes a simple two-category label less useful, but it does not prove that an assistant meets any particular AGI definition. Google DeepMind’s proposed AGI framework treats performance, generality, and autonomy as separate dimensions to evaluate. Our generative AI guide explains what content generation adds without treating it as a synonym for AGI.
How to evaluate a narrow AI tool
Start with the exact job you want the system to perform, not its marketing category. Ask:
- What output should it produce: a classification, recommendation, draft, or action?
- Was it tested on examples like your own data, including unusual cases and costly errors?
- Where is the boundary of its intended use, and when should a person take over?
- How will you monitor errors and changes in the input data after deployment?
For the hypothetical contract classifier, a firm could measure missed clauses and false flags on a representative sample, then require a lawyer to check the original document. A good score on ordinary templates would not establish performance on unfamiliar drafting or different contract types.
NIST’s AI Risk Management Framework recommends evaluating performance under conditions similar to the intended deployment, documenting where results may not generalize, and monitoring the system in use. The framework is voluntary guidance, not a promise that a tested system cannot fail. Our responsible AI guide covers broader oversight questions.
The practical lesson is to identify what an AI system has actually been shown to do, then test and supervise it for the work you need. A powerful specialized result is valuable. It is not evidence of general competence in every other setting.
Sources and date
Definitions and examples were checked September 25, 2026: OECD’s AI-system definition, Google on Gmail spam filtering, Google DeepMind on AlphaGo, Google DeepMind’s AGI framework, and NIST on evaluation in the intended context. Product implementations and research definitions can change.