The perfect prompt isn’t enough anymore

"The Perfect Prompt" was the first AI skill, not the last.

Step 1: The AI responds

By the end of 2022, all you’ll need to do is type a question or instruction into a chat window. You’ll receive a text, an explanation, or an idea in return. The person formulates the task, the AI generates a suggestion, and the person decides what to do with it.

The misconception at this stage persists to this day: whoever knows the perfect prompt has mastered AI. In reality, however, what you’re really learning here is simply how to formulate your request clearly enough so that a language model can generate useful output from it.

That’s useful. But it doesn’t yet constitute comprehensive AI expertise.

Even at this initial stage, it is important to understand that a language model does not retrieve the truth, nor does it engage in dialogue in the social science sense. It processes input, calculates statistically probable continuations, and can formulate responses convincingly without knowing whether the result is true, complete, or appropriate.

The corresponding human skill is therefore not merely “prompting,” but rather a fundamental ability to evaluate: Which parts are correct? What is missing? Which source supports the statement? What do I need to judge for myself?

Level 2: AI processes more than just text

Starting in 2023, and increasingly in 2024, the text-generating model will evolve into a multimodal system. AI will process images, language, audio, video, code, spreadsheets, and large document collections. At the same time, ever-better systems for generating images, voices, and moving images are emerging. For example, Runway developed consistent generative video worlds, Google combined different media formats in Gemini, and Meta released multimodal models with open weights. Runway: Gen-4, Google: Gemini 2.0, Meta: Llama 4

This doesn’t just change the amount of material that can be produced; it also broadens the scope of perception. An AI can analyze a screen, explain a graphic, transcribe a conversation, examine a video, or compare different documents with one another.

Those who now simply enter data and accept the results may produce more, but they can easily lose track of where statements, images, and interpretations actually come from.

That is why we need more expertise, not less. Only those who understand a topic themselves can recognize where AI sounds plausible but is still off the mark. Added to this are data and media literacy: What materials is the system allowed to process? What has been altered? Which source is still recognizable? Is an image documentary or synthetic? Can I trace how a result was generated?

The key shift is this: It is no longer enough to simply request good content. People must curate information spaces.

There is no longer “the AI”

Alongside technological advancements, the market has become more diverse. In addition to the large, closed platforms from OpenAI, Google, and Anthropic, open or openly available model families such as Llama, Mistral, and DeepSeek have emerged. In early 2025, DeepSeek-R1 demonstrated that high-performance reasoning models do not have to come exclusively from the major American platform companies. Meta brought smaller models to mobile devices; Apple embedded AI processing directly on devices and within a highly secure cloud infrastructure. DeepSeek-R1, Meta: Llama 3.2 for mobile devices, Apple: On-Device Processing and Private Cloud Compute

This is not a minor technical issue. It changes the question of responsibility. People and organizations no longer need to simply learn how to operate a system. They must choose which system is suitable for which purpose, where their data will be processed, what dependencies arise, and whether a more affordable, local, or open model is sufficient for the task.

In addition to operational competence, there is also selection competence and decision-making competence.

Stage 3: AI Becomes Invisible in the Workplace

In the next phase, AI will be less visible as a standalone chat window. It will be embedded in search engines, office programs, development environments, creative software, cell phones, and corporate knowledge systems. It will summarize, sort, supplement, prioritize, and make suggestions without users having to consciously type a prompt for every action.

This changes the nature of how the system is used. With a visible chat window, you at least know that an AI system is involved. With embedded AI, automated preselection quickly becomes a natural part of the work environment.

The key competency here is contextual and systemic competence: In what ways is AI actually involved here? What data does it receive? What criteria does it use to sort the data? Which suggestions are displayed and which are not? What happens automatically before a human even makes a decision?

Blog-Abo

Neue Beiträge direkt im Postfach.

Blog abonnieren

Even the perfect prompt is of little help at this point. The key factors often lie not in the individual voice command, but in the system settings, data access, stored rules, and organizational routines.

Level 4: The AI processes multi-step tasks

Since 2025, powerful systems have been able to track a task across multiple steps. They research, compare sources, analyze files, write and test code, select tools, and correct interim results. As early as late 2024, Google explicitly referred to Gemini 2.0 as a model for an “agent-driven era.” Anthropic describes agents as systems that autonomously combine tools, access to information, and feedback loops. Google: Gemini 2.0, Anthropic: Building Effective Agents

You no longer just enter a single command. You set a processing workflow in motion.

This requires a different skill set than before: it’s no longer just about asking good questions, but about developing good tasks and evaluation criteria. What exactly should be achieved in the end? Which sources are acceptable? Which steps must be documented? When is a result good enough? When must the system stop or request a human decision?

Anyone who can’t answer these questions isn’t in control of the AI process. They’re just hoping that something useful will come out of it in the end.

This is where judgment becomes a technical prerequisite for the first time. Before a system can work toward a goal on its own, someone must have determined what constitutes success, failure, or overstepping a boundary.

In February 2026, it became clear just how serious this shift was. Anthropic released Claude Cowork, which made development tools accessible to a broader professional audience. Within a few weeks, software companies such as Thomson Reuters, Salesforce, and Adobe collectively lost about $285 billion in market value because licensing models for individual human workstations were suddenly called into question. Anthropic engineer Boris Cherny described his daily work at that time as no longer prompting himself, but rather building systems that prompt for him: “My job is to write loops.” That is the short formula for Level 4: no longer giving individual instructions, but designing systems that know for themselves when they are done.

Level 5: The AI acts with delegated authority

Currently, agent-based systems have access to browsers, files, calendars, accounts, and software. They no longer merely process information; they can also modify entries, create files, operate applications, and trigger actions. For example, since 2024, Anthropic has enabled computer control via the screen, mouse, and keyboard; in 2025, advanced features for the independent selection and use of tools were added. Anthropic: Computer Use, Anthropic: Advanced Tool Use

Just how closely this delegation is tied to real-world applications is demonstrated by the OpenClaw agent developed by Austrian developer Peter Steinberger. The tool sends emails, coordinates appointments, checks in during trips, and conducts research without requiring anyone to approve every step. In March 2026, thousands of people lined up in Shenzhen to have the free, open-source program installed. Five months later, the ambivalence surrounding the tool has become apparent: Chinese authorities banned state-owned enterprises from using it because its extensive system privileges could lead to data leaks and unintended deletions, while local governments in Shenzhen and Wuxi are simultaneously promoting companies that build on the same tool.

In this context, “autonomous” does not mean that it has become an independent entity. It means, first and foremost, that between what a person instructs the system to do and what ultimately happens, there are many small, automated decisions that are no longer approved on a case-by-case basis.

The rights remain on loan. People and organizations decide which data, tools, and courses of action a system is allowed to access. They remain responsible even when they no longer carry out every intermediate step themselves.

This requires the ability to delegate and monitor. What tasks can the system handle? What permissions are necessary? What must remain reversible? Which actions require approval? Where are protocols, controls, and a clear escalation process needed?

What’s interesting here is that experienced users don’t simply grant agents more and more freedom. According to an Anthropic analysis of real human-agent interactions, as experience increases, so do both the level of autonomy granted and the frequency of targeted interventions. Competent use, therefore, does not mean less control, but a different kind of control: less approval of every action, and more deliberate intervention at critical points. Anthropic: Measuring AI Agent Autonomy in Practice

So far, this trend has affected only a limited portion of the workforce. In the 2026 German Social Collaboration Study, fewer than one-third of the organizations surveyed were using AI agents at all. Their use was concentrated in IT support, customer service, and project management; it was rare in production and actual value creation. However, the study is based on a limited survey of executives and does not allow for generalizations about all German companies.

Nevertheless, the technology is more advanced than many of the organizations that are expected to use it in the future. Those who put controlled delegation to the test today are building a capability whose importance is likely to grow as agent-based systems become more widespread.

Stage 6: AI Becomes Infrastructure

The next likely shift isn’t simply an even more powerful chatbot. AI is becoming the background infrastructure of organizations. Multiple systems will handle different subtasks, access shared databases, and pass results to one another. Employees will then no longer work solely with a single AI application, but within a technical and organizational environment in which human and machine contributions are intertwined.

This is a plausible scenario, not a definitive prediction. It remains to be seen how reliable such systems will be, how quickly companies will adapt their data and processes for them, and what regulatory constraints will arise.

However, the question of competence is already shifting from the individual user to the organization. It then becomes a matter of roles, responsibilities, data access, quality standards, learning paths, and power. Who is authorized to set goals? Who monitors results? Who is liable for errors? Which human skills must be preserved? Which tasks can be eliminated, and which ones absolutely must not disappear?

At this stage, individual AI expertise is generally no longer sufficient. What is needed is organizational system expertise.

The hook that stays attached the whole time

This is the part that most narratives of progress overlook. Each stage requires more judgment than the previous one. But judgment isn’t developed on the drawing board. It arises precisely from the work involved in the lower stages: writing a first draft yourself, conducting your own research, finding your own mistakes, and experiencing a difficult situation multiple times.

If this work is delegated to AI from the very beginning, there may later be a lack of experience needed to even evaluate an AI-generated result.

This affects, first and foremost, those who are just starting out. Young professionals who never conduct their own research—because AI already does it for them—will not easily learn to recognize poor research. Those who no longer develop their own first drafts have fewer opportunities to understand why a text works or doesn’t. Those who have never had to look for a simple mistake will have a harder time identifying a complex one.

Anyone who automates entry-level jobs today is therefore also determining the number of people who, ten years from now, will still be capable of professionally reviewing automated results—most often without even treating it as a decision.

That is the true paradox of judgment: The very technology that makes human judgment more valuable can erode the pathways of experience through which that judgment has developed thus far.

What this means for assessing one’s own position

The obvious reaction to these stages would be to place oneself somewhere along the spectrum and stay there. That would be exactly the wrong conclusion.

A person is not entirely at “Level Three.” While drafting text, they may already use agent-based workflows, but when conducting research, they may still accept individual answers without verifying them. A company may use agents for IT support while organizing its HR processes entirely manually. The levels describe tasks, not people.

The key question, therefore, is not just: Do I use AI?

The question is: Does my understanding of AI expertise still align with what my specific tasks now require?

Anyone who practices perfect prompts at Level One while their own tasks have long since required Level Four is working very hard on the wrong skill. Conversely, anyone who delegates tasks immediately without having mastered the lower levels themselves is relying on judgment that may never have developed.

Prompting was the first visible AI competency. It is followed by subject matter expertise, source criticism, context creation, system selection, process design, delegation, monitoring, and organizational responsibility.

The perfect prompt is no longer enough. But even the perfect agent won’t solve the problem. The scarce resource remains the human being who can make reasoned decisions about what the system should do, how to recognize a good result, and what tasks should not be delegated to machines.

Source verification by Claude (August 22, 2026)

All technical citations/data verified against primary sources: Runway Gen-4 (March 31, 2025), Meta Llama 4 (April 5, 2025), DeepSeek-R1 (January 20, 2025), Meta Llama 3.2 (September 25, 2024), Apple On-Device/Private Cloud Compute (WWDC 2024), Google Gemini 2.0 “agentic era” (December 11, 2024), Anthropic Computer Use (October 22, 2024), Anthropic Advanced Tool Use (November 24, 2025), Anthropic “Measuring AI Agent Autonomy” (Auto-Approval ~20→40%+, Interrupt Rate ~5→9%), German Social Collaboration Study 2026 (TU Darmstadt/Buxmann + Campana & Schott, >200 executives), OpenClaw/Steinberger (March 2026 Shenzhen hype; since then, a partial ban for state-owned enterprises alongside support from local governments), Claude Cowork/SaaSpocalypse (released in late January 2026, ~$285 billion in market value lost).

FLOWCAMPUS · KI-Orientierung im Austausch
Wo passt KI in deine Arbeit?

Andere Use Cases helfen nur bedingt. In KIOA arbeitest du an deinem eigenen: Du klärst, was sich für dich zu verändern lohnt, probierst es zwei Wochen im Arbeitsalltag aus und entscheidest danach, wie du weitergehst.

Start 18.9. Berlin oder hybrid 2 Wochen Praxis 299 € zzgl. MwSt.
Programm ansehen →
Anmeldung bis 11. September · 6 bis 12 Personen
Scroll to Top