AI agent
Autonomous artificial intelligence agent
From Wikipedia, the free encyclopedia
An AI agent is an artificial intelligence system that can pursue goals, use software or other tools, and take actions with some level of autonomy.[1][2][3] Agentic AI is a related and inconsistently defined term for systems designed to exhibit such goal-directed autonomy, sometimes through multiple coordinated agents.[4]
Overview
There is no single standard definition of what an AI agent is.[4][5][6] Common attributes of AI agents include goal-directed behavior, natural language interfaces, the capacity to use external tools, and the ability to perform multi-step tasks. Their control flow is frequently driven by large language models (LLMs). Agent systems may also include memory components, planning logic, tool interfaces, and orchestration software for coordinating agent components.[3][6]
A common application of AI agents is task automation: for example, booking travel plans based on a user's prompted request.[7][8]
Companies such as Google, Microsoft and Amazon Web Services have offered platforms for building and deploying AI agents.[9][10][11] Open protocols have been proposed for connecting agents to tools and enabling communication between agents, including Model Context Protocol (MCP) and Agent2Agent (A2A).[12][13]
In December 2025, Linux Foundation announced the formation of the Agentic AI Foundation (AAIF), with the goal of ensuring that agentic AI evolves transparently and collaboratively.[14][15]
History
The concept of an intelligent artificial agent predates current LLM-based systems. Oliver Selfridge's 1959 "Pandemonium: A Paradigm for Learning" described a system of interacting computational units for pattern recognition.[16] During the 1990s, agent research developed through work on agent-oriented programming, the belief–desire–intention software model (BDI), and formal accounts of intelligent-agent architectures.[17][18][19]
Research on LLM-based autonomous agents expanded in the early 2020s as language models were combined with planning, memory, and external tools.[3] OpenAI introduced function calling in its API in June 2023, allowing models to return structured requests for software functions.[20] In November 2024, Anthropic introduced the Model Context Protocol (MCP), an open protocol for connecting AI assistants to external data sources and tools.[21] The term "agentic" began to be used with greater frequency in 2024 and was popularized in part by researcher Andrew Ng.[22]
Training and testing
Researchers train and evaluate AI agents in interactive environments that expose them to sequences of observations, actions, and feedback.[23] Video-game environments such as Minecraft and No Man's Sky and benchmark environments that reproduce common websites have been used for this purpose.[24][25]
Autonomous capabilities
The Financial Times compared AI agents' autonomy to the SAE classification of self-driving cars, likening most applications to level 2 or level 3, with some achieving level 4 in highly specialized circumstances and level 5 being theoretical.[26]
Cognitive architecture
Ken Huang proposed an AI agent reference architecture that consists of seven interconnected layers, with each layer building on the functionality of the layers beneath it:[27]
- Layer 1: Foundation models – provide the core models that power the agent.
- Layer 2: Data operations – manages the data infrastructure required for AI agent operations, including vector databases, data loaders, and RAG.
- Layer 3: Agent frameworks – software that manages the AI agents.
- Layer 4: Deployment and infrastructure – the technical foundation of the AI agents.
- Layer 5: Evaluation and observability – the safety and performance of AI agents.
- Layer 6: Security and compliance – a protective framework for safe operation and compliance with regulatory boundaries. At this layer, security and compliance features embedded into all the AI agent stack layers are integrated together.
- Layer 7: Agent ecosystem – represents the AI agents' interface with real-world applications and users.
Agent harness
In developer documentation, an agent harness refers to the software layer surrounding a large language model that enables it to function as an AI agent. It commonly manages prompts, context, tool use, memory, execution state, operational constraints, sandboxes, permissions, and the processing of results. The harness connects the model to internal and external computer hardware, software, data files, databases, web browsers, command-line interfaces, and application programming interfaces, while controlling how the agent accesses and uses these resources to complete multi-step tasks.[28][29]
Orchestration patterns
Autonomous agents are often integrated with other agents or specialized tools to execute complex tasks. These configurations, known as orchestration patterns or workflows, include the following:[30]
- Prompt chaining: A sequence where the output of one step serves as the input for the next.
- Routing: Directing an input to a specialized downstream task or tool.
- Parallelization: The simultaneous execution of multiple tasks.
- Sequential processing: A fixed, linear progression of tasks through a predefined pipeline.
- Planner-critic: An iterative pattern where one agent generates a proposal and another evaluates it to provide feedback for refinement.
Multimodal AI agents
In addition to large language models (LLMs), vision-language models (VLMs) and multimodal foundation models can be used as the basis for agents.[31] Allen Institute for AI released the open-weight Molmo family of vision-language models in 2024.[32] Nvidia released a framework for developers to use VLMs, LLMs and retrieval-augmented generation for building AI agents that can analyze images and videos, including video search and video summarization.[33][34] Microsoft released a multimodal agent model – trained on images, video, software user interface interactions, and robotics data – that the company claimed can manipulate software and robots.[35]
Applications
As of April 2025, per the Associated Press, there are few real-world applications of AI agents.[36]
The Information divided AI agents into seven archetypes: business-task agents, for acting within enterprise software; conversational agents, which act as chatbots for customer support; research agents, for querying and analyzing information (such as OpenAI Deep Research); analytics agents, for analyzing data to create reports; software developer or coding agents (such as Cursor); domain-specific agents, which include specific subject matter knowledge; and web browser agents (such as OpenAI Operator).[37]
By mid-2025, AI agents were being used in video game development,[38] gambling (including sports betting),[39] cryptocurrency wallets[39] (including cryptocurrency trading and meme coins[40]) In August 2025, New York Magazine described software development as the most definitive use of AI agents.[41] The Information noted AI coding agents and customer support as the primary uses of AI by businesses by October 2025, although a decline in the expectations of AI capabilities was also noted.[42]
In November 2025, The Wall Street Journal reported that few companies that deployed AI agents have received a return on investment.[43]
Applications in government
Several government bodies in the United States and United Kingdom have deployed or announced the deployment of agents at the local and national level. The city of Kyle, Texas deployed an AI agent from Salesforce in March 2025 for 311 customer service.[44] In November 2025, the Internal Revenue Service stated that it would deploy Salesforce AI agents for the Office of Chief Counsel, Taxpayer Advocate Services, and the Office of Appeals.[45] That same month, Staffordshire Police announced that they would trial Agentforce agents for handling non-emergency 101 calls in the United Kingdom starting in 2026.[46] In December 2025, the Department of Neighborhoods in Detroit, Michigan, partnered with a local business to deploy a customer service AI agent in two city districts.[47]
In February 2025, Thomas Shedd, the director of the Technology Transformation Services, proposed using AI coding agents across the United States federal government.[48] In April 2025, a recruiter for the Department of Government Efficiency proposed using AI agents to automate the work of about 70,000 United States federal government employees as part of a startup with funding from OpenAI and a partnership agreement with Palantir. This proposal was criticized by experts for its impracticality and the lack of corresponding widespread adoption by businesses.[49]
In December 2025, the Food and Drug Administration announced that it would offer "agentic AI capabilities" to its staff for "meeting management, pre-market reviews, review validation, post-market surveillance, inspections and compliance and administrative functions."[50] That same month, the United States Department of Defense launched GenAI.mil, an internal platform for American military personnel to use generative AI-based applications based on Google Gemini, including "intelligent agentic workflows". Defense Secretary Pete Hegseth listed applications such as "[conducting] deep research, [formatting] documents and even [analyzing] video or imagery at unprecedented speed."[51] In December 2025, the United States Immigration and Customs Enforcement agency signed a contract with a company for its Enforcement and Removal Operations department to use AI agents for skip tracing.[52]
Operating systems
In November 2025, Microsoft released a test software build of Windows 11 that included agents to run background tasks, with the ability to read and write personal files.[53] In December 2025, ByteDance released Doubao, an AI agent that can be integrated into smartphone operating systems, particularly the Nubia M153 by ZTE.[54] Several apps in China blocked or restricted the agent, citing privacy and security concerns, including WeChat,[54] Alipay, Taobao, Pinduoduo, Ele.me,[55] and local banks.[56]
Web browsing
Web browsers with integrated AI agents are sometimes called agentic browsers. Such agents can perform tasks and browser actions on the user's behalf.[57]
In 2025, Microsoft launched NLWeb, an agentic web search replacement that would allow websites to use agents to query content from websites by using RSS-like interfaces that allow for the lookup and semantic retrieval of content. Within weeks of release, NLWeb was found to have created security issues and expose information about users to third-party servers.[58]
Proposed benefits
AI agents have been proposed as a means of increasing personal and economic productivity,[7][59] fostering greater innovation,[60] and liberating users from monotonous tasks.[60][61] However, Parmy Olson's Bloomberg opinion piece argued that agents are best suited for narrow, repetitive tasks with low risk.[62] Conversely, researchers suggest that agents could be applied to web accessibility for people with disabilities,[63] and researchers at Hugging Face propose that agents could be used for coordinating resources such as during disaster response.[64] The R&D Advisory Team of the BBC views AI agents as being most useful when their assigned goal is uncertain.[65] Erik Brynjolfsson suggests that AI agents are more valuable in enhancing, rather than replacing, humans.[66]
Concerns
Research on increasingly agentic systems has identified risks involving weakened human oversight, privacy, bias, manipulation, misinformation, and systemic or long-range harms.[67][6] Their autonomy can also complicate monitoring, accountability, and the allocation of legal responsibility.[68] Agents that use external tools may be vulnerable to prompt-injection attacks in which untrusted data causes the agent to perform unintended actions.[69] Enterprise deployment has additionally raised contracting concerns related to liability allocation, data ownership, and legal accountability.[70]
AI agents may also have a negative impact on the environment due to high energy usage.[65][71][72] Nvidia CEO Jensen Huang predicted that AI agents would require 100 times more computing power than LLMs.[73] There is also the risk of increased political corruption, as AI agents may not question instructions in the same way that humans would.[74]
Journalists have described AI agents as part of a push by Big Tech companies to "automate everything".[75] Several of those companies' CEOs stated in early 2025 that they expect AI agents to eventually "join the workforce".[76][77] In TheAgentCompany, a workplace benchmark accepted at NeurIPS 2025, the best evaluated agent autonomously completed 30.3% of the tasks.[78] Other researchers had similar findings with Devin AI[79] and other agents in both formal business settings[80] and freelance work.[81]
In June 2025, CNN argued that CEOs's statements on AI replacing their employees were a strategy to "[keep] workers working by making them afraid of losing their jobs."[82] Tech companies have pressured employees to use generative AI models in their work, including AI coding agents. Brian Armstrong, the CEO of Coinbase, fired several employees who did not.[83][84] Some business leaders have replaced some of their employees with agents, but have said that the agents would need more supervision than those employees.[42]
Large technology companies such as Salesforce, Klarna and IBM announced layoffs in 2025, replacing hundreds of their employees in human resources or customer service with AI agents.[85][86][87] However, Klarna later rehired several human employees.[86]
Financial authorities have warned that more complex and autonomous "agentic" AI could become a channel for systemic risk in finance.[88] They distinguish these systems from other AI because they can pursue goals over many steps, call tools, and carry out tasks with relatively little human intervention. In workshops with regulators, central‑bank officials, and industry specialists, participants highlighted risks both from agentic systems built inside financial institutions and from tools offered by technology firms that can initiate or execute financial actions. In one 2025 forum, 44% of experts surveyed judged autonomous or agentic AI systems to be the most likely current source of AI‑related systemic risk in finance.[88]
In March 2025, Scale AI signed a contract with the United States Department of Defense to work with them, in collaboration with Anduril Industries and Microsoft, to develop and deploy AI agents for the purpose of assisting the military with "operational decision-making".[89] In July 2025, Fox Business reported that the company EdgeRunner AI built an offline agent, compressed and fine-tuned on military information, with the CEO seeing more common LLMs as "heavily politicized to the left". As of that time, the company model is being used by the United States Special Operations Command in an overseas deployment.[90] A study of LLM agents in simulated wargames found escalatory and difficult-to-predict behavior, leading its authors to recommend caution before using autonomous agents in strategic military or diplomatic decision-making.[91]
NY Mag unfavorably compared the user workflow of agent-based web browsers to Amazon Alexa, which was "software talking to software, not humans talking to software pretending to be humans to use software."[92] The same outlet described web browser agents and computer-use agents as an attempt to "click-farm the entire economy."[93]
Agents may get stuck in infinite loops.[94][95]
Since many inter-agent protocols are being developed by large technology companies, there are concerns that those companies could use these protocols to benefit themselves.[96]
In June 2025, Gartner described the rebranding of existing assistants, chatbots, and robotic-process-automation products as agents without substantial agentic capabilities as "agent washing".[97]
Researchers have warned about the impact of providing AI agents access to cryptocurrency and smart contracts.[40]
During a vibe coding experiment, a coding agent by Replit deleted a production database during a code freeze, "[covered] up bugs and issues by creating fake data [and] fake reports" and responded with false information.[98][99]
OpenAI co-founder Andrej Karpathy criticized AI agents as being ineffective and promoting AI slop.[100]
Issues with multi-agent systems include coordination and communication between component agents, inconsistent performance, and challenges in debugging.[101]
In November 2025, Anthropic claimed that a group of hackers sponsored by China attempted a cyberattack against at least 30 organizations by using Claude Code in an agentic workflow, and that several of these infiltrations had succeeded.[102] However, independent cybersecurity researchers questioned the significance of Anthropic's findings.[102][103]
Whittaker argued that the push by Big Tech companies to deploy AI agents risked security vulnerabilities across the Internet.[104]
Agentic misalignment
"Agentic misalignment" refers to situations in which an AI agent pursues strategies that conflict with its designers' intentions. In 2025, Anthropic reported controlled simulations in which models sometimes engaged in harmful actions, such as blackmail or corporate espionage, when those actions appeared necessary to achieve an assigned goal or avoid replacement. Anthropic emphasized that the scenarios were deliberately constructed stress tests and that it was not aware of such behavior occurring in real-world deployments.[105]
Security
Threat modeling frameworks
Threat-modeling approaches applied to AI agents include general software-security models and AI-specific frameworks:
- STRIDE: A Microsoft model that identifies threats across six categories: spoofing, tampering, repudiation, information disclosure, denial of service, and elevation of privilege.[106]
- MITRE ATLAS: A knowledge base of adversary tactics and techniques for AI systems.[107]
- OWASP GenAI Security Project: Provides guidance on vulnerabilities associated with generative AI and large language model integration, including guidelines specifically for agentic applications.[108]
- MAESTRO: A framework from the Cloud Security Alliance for assessing risks in multi-agent AI systems throughout their lifecycle.[109]