Opinions on how intelligent artificial intelligence is vary widely. While the latest models now not only master common school-leaving exams but already pass entrance tests at renowned universities, the various AI services à la ChatGPT, MS Copilot, or Gemini are still limited in logic, autonomy, and decision-making ability—and thus in independent action. That could soon change, however. The development of AI systems is currently transforming at breakneck speed from task-specific, static AI agents that require human input (prompts) to interactive and self-learning agents. Thanks to a higher level of autonomy, these can adapt to their environment and, with their growing maturity, move another step closer to “Artificial General Intelligence” (AGI). According to OpenAI founder Sam Altman, AGI is capable of performing economically valuable work as creatively, intelligently, and flexibly as a human—if not even better.
Task-specific AI vs. generalist agents
Task-specific AI is already widely used today in the financial and insurance industry. It serves to automate simple, structured tasks; most chatbots in use are a good example. They are fed with knowledge about data and workflows and operate within the predefined framework, which allows them to relieve valuable specialists and increase efficiency—however, they decide only according to a preprogrammed scheme and can adapt to new situations only within narrow limits.
Agentic AI is developing toward generalist agents that possess new capabilities, greatly decoupling their learning, decision-making, and actions from human intervention. They are based on multimodal systems that efficiently process a large amount of very diverse (environmental) data and can derive from it a broad range of agent-based multimodal interactions. The data can be text, but also video or robotics sequences (hence: multimodal).
A high degree of autonomy
In general, it is inherent to AI that it far surpasses humans in processing large amounts of data in the shortest time. New AI models enable AI agents with a higher degree of autonomy, so that they do not have to constantly interact with humans. Work steps are planned and carried out independently.
An AI’s performance is measured by the “General AI Assistant (GAIA)” benchmark. At the lowest level are AI models that are capable of completing tasks that require no tool or at most one tool and can be solved in a maximum of five steps. At Level 2, AIs stand out on tasks for which they need five to ten steps and must combine several tools. Level 3, finally, presupposes a general understanding of the world in order to tackle tasks for which the number of work steps and tools is unlimited. The latest AI models are already coming quite close to human intelligence, achieving over 80 percent at Level 1 (vs. humans: 100), 70 percent at Level 2 (vs. humans: 92), and almost 60 percent at Level 3 (humans: 87).