For most people, AI still means a chat window that writes text. The more consequential development is quieter: systems that can take in screenshots, documents, voice notes and dashboards together, work out what needs to happen, and then carry out the steps across your software. Multi-modal perception combined with agentic action turns AI from a tool you consult into a worker you delegate to.
That shift is a management question before it is a technology question. When software can own a process end to end, leaders have to decide which processes to hand over, who is accountable for the results, and how the organization changes around it. The companies that treat agents as a new kind of capacity to be managed, rather than a feature to be switched on, are the ones pulling ahead.
Why multi-modal plus agentic is a step change
Each capability matters on its own. Multi-modal models remove the need to convert everything into text before AI can use it: a support agent can look at the customer’s photo of an error, listen to the voicemail and read the account history in one pass. Agentic systems plan and act: they choose tools, call APIs, update records and check their own work.
Together they close the loop. Earlier AI could summarize the situation and leave a person to act. Now the same system can notice the problem in a dashboard image, find the related tickets, draft the fix and file the change request. The practical effect for a leader is that work which previously required a person to stitch together four systems can be delegated as a single unit.
McKinsey’s State of AI survey, published in November 2025, shows where most organizations stand. Some 88% of respondents said their organizations use AI in at least one business function, and 62% were at least experimenting with AI agents. But only 23% were scaling an agentic system anywhere in the enterprise, and only about 6% qualified as high performers seeing significant bottom-line impact. The gap between experimenting and benefiting is mostly a leadership and operating-model gap.
Shift the question from “what can AI do” to “what can AI own”
Asking what AI can do produces long lists of demos. Asking what AI can own produces decisions. Ownership means the agent is responsible for a defined outcome, within limits, with a person accountable for its performance, much as you would delegate to a new team member.
Processes that are good candidates for AI ownership tend to share traits:
- Clear inputs and a definable “done.” Resolving a password reset, reconciling a vendor statement, or preparing a weekly competitive brief.
- Mixed media that slows people down. Work that requires reading PDFs, looking at images and listening to calls is where multi-modal agents save the most time.
- Reversible actions. Mistakes can be caught and corrected before they cause lasting harm.
- Measurable quality. You can sample the output and judge whether it was right.
Processes that involve irreversible financial commitments, legal judgments, safety decisions or sensitive personnel matters should keep a person as the decision-maker, with the agent preparing the work.
Where leaders are delegating first
The early patterns are consistent across industries:
- Customer support. Agents that read the ticket, the attached screenshot and the account history, resolve routine issues and route the rest with a summary.
- Revenue operations. Agents that analyze recorded sales calls, update the CRM and draft follow-ups for the rep to approve.
- Marketing production. Pipelines that plan, draft, create images and schedule content, with a human editor approving before publication.
- Market and product research. Agents that monitor competitors’ sites, pricing and releases and produce a daily or weekly brief.
- Back-office operations. Agents that coordinate tasks across SaaS platforms, such as onboarding a new customer across billing, support and project tools.
For the day-to-day operational side of multi-modal AI in smaller firms, see Multimodal AI in 2026.
A leader’s playbook for delegating to agents
- Name an accountable owner for every agent. Someone on the business side answers for its results, just as they would for a team member’s.
- Write the job description. Define the outcome, the systems the agent may touch, the actions it may take on its own and the ones that need approval.
- Redesign the workflow, not just the task. McKinsey’s research found high performers were far more likely to have fundamentally redesigned workflows. Dropping an agent into an unchanged process usually moves the bottleneck rather than removing it.
- Set limits and log everything. Least-privilege access, spending caps and a complete record of what the agent saw and did.
- Measure like a manager. Track volume, quality from regular sampling, escalation rate and cost per outcome, and review them monthly.
- Plan for the people. Decide how freed-up time will be used, retrain staff to supervise and improve agents, and communicate early.
The risks leaders own
Delegation does not transfer accountability. An agent that can read email and act on it can be manipulated by instructions hidden in the content it reads. An agent with broad permissions can make a small error at scale. Multi-modal inputs also mean more sensitive data, such as recorded calls and images of people, flowing through AI systems. These are governance issues for the leadership team, not only the IT department; our overview of AI agent security covers the controls in detail.
There is also a quieter risk: over-trusting fluent output. Agents present their work confidently whether or not it is correct, so sampling results, even after the pilot phase, is part of the manager’s job. Leaders who set the expectation that agent output is reviewed like any junior colleague’s work avoid the slow drift from delegation into abdication.
Frequently asked questions
What is a multi-modal AI agent?
It is an AI system that can take in several kinds of input, such as text, images, audio and documents, and then plan and carry out actions in other software to reach a goal, rather than only producing an answer.
How should leaders decide which processes to hand to AI agents?
Start with processes that have clear inputs, a measurable definition of done, reversible actions and a lot of mixed-media handling. Keep people as decision-makers where actions are irreversible, legally significant or safety-related.
Do AI agents replace employees?
In most early deployments they absorb the repetitive portion of a role, while people handle exceptions, relationships and judgment. The organizations that benefit most redesign roles deliberately rather than leaving the change to chance.
Lead the shift deliberately
Delana Technologies helps leadership teams decide what AI should own, design the workflows and guardrails around it, and deploy multi-modal agents securely. Explore our AI consulting and agentic AI solutions, call 239.414.5126 or contact us.
Sources: McKinsey & Company, “The state of AI in 2025: Agents, innovation, and transformation” (November 2025); IT Brief coverage of the McKinsey survey.
