Signal
Adopting an AI artifact requires evidence, validation, and human decisions to pass to the next person or agent
When an AI agent is asked to research, reconcile, or prepare a document, it can return a result quickly.
That report does not finish the work.
Without knowing which conditions were met, which materials were used, where the work stopped, and what a person must decide, the artifact cannot safely move into the business process.
AI assistants answer questions, while AI agents use tools to execute work.
Agentic systems combine multiple steps or agents, builders create them, and governance platforms manage permissions and usage.
Enterprise work continues with goal setting, changing conditions, artifact integration, evidence review, and approval.
OpenAI Workspace Agents run long-running work within organization-defined data and tool permissions, request approval for selected operations, and expose configuration, updates, and runs to administrators.
Microsoft and Google evaluators separately inspect task completion, tool selection, tool inputs, and use of tool outputs. NIST Evaluation Probes compare agent claims with a human-curated trusted corpus and accumulate the evidence mapping in a machine-readable audit trail.
In a customer statement published by OpenAI, Rippling reports that its sales workspace agent researches accounts, summarizes Gong calls, and posts deal briefs to Slack, turning work that had taken representatives five to six hours per week into a process that runs automatically on every deal.
This is a deployed example of an agent operating across tools and a shared team process. The example alone does not establish that human business decisions are promoted into task contracts or evaluators for separate future work.
Taken together, these sources point to a work environment that preserves who executed each step, under which authority, how it was validated, and which decisions a person accepted after work was delegated to AI.
Definition
This briefing calls the environment that carries objectives, constraints, execution, evidence, validation, and human decisions a HACW
Solving this problem requires an environment that preserves the request's objective and constraints, execution record, evidence, validation results, and human decisions when the business owner or AI agent changes.
HACW is a working hypothesis derived from public evidence, not a detailed standard that Gartner has publicly established.
Gartner's public abstract for the 2026 Hype Cycle for Digital Workplace Applications says digital-workplace priorities have shifted toward governing embedded AI and preparing for “human-agent work models.”
Here, workspace does not mean only a chat interface or a shared folder.
It means a shared record that preserves the same objective, current owner, permitted actions, completed and unfinished items, evidence, and validation results when an agent changes or the work returns to a person.
Working definition
A Human-Agent Collaboration Workspace (HACW) is a work environment that records the objective and prohibited actions, owners and authority, agent execution, artifacts and sources, task-specific validation, and human decisions for work requested by a person. The record remains available when the responsible person or AI agent changes.
Human decisions are stored with reasoning, applicability, exceptions, and expiry. Only decisions that can be tested under the same conditions become task contracts, evaluators, or authorization policies that future tasks may use without another human review.
Shared work record across every step
Records
Goal, owner, authority, progress, artifacts, evidence, validation results, approvals, and unresolved issues
When an agent or person changes, everyone continues from the same objective, progress, permitted actions, artifacts, validation results, and unresolved issues.
AI assistant
Uses conversation as the entry point and returns an answer, summary, or draft. A person carries the result into the next task.
AI agent and agentic AI
Uses tools to execute one or more steps toward a goal. Execution is the central object.
Build, runtime, and governance platforms
Create, connect, run, and monitor agents. They provide essential developer and administrator functions, but do not necessarily keep the human request, execution, validation, decision, and future rule update in a shared record.
HACW
Shares requests, owners, permitted actions, artifacts, evidence, validation results, decision reasoning, active business rules, and unresolved issues. Its central question is who executed what and which decisions may become rules that future tasks use without another review.
Microsoft Research has described collaboration through mutual goal understanding, proactive task co-management, and shared progress tracking.
Its Interaction, Process, Infrastructure framework also makes the structure of activities explicit and inspectable.
Together, these implementations and research support a HACW condition: the process must not disappear behind a conversation. Owners, authority, progress, validation results, and human decisions remain in the same work record as the artifact.
A HACW keeps these records in a shared record that the business owner, orchestrator, executor, and evaluator can inspect rather than hiding them inside a single agent.
Operating Model
Different authority and judgment criteria separate business owners, orchestrators, specialist agents, and evaluators
Setting objectives and prohibited actions, assigning tasks, updating data, and validating artifacts require different authority and judgment criteria.
In a HACW, a business owner therefore sets the conditions of the work, an orchestrator selects owners, and specialist agents or existing systems execute it.
Task-specific evaluators inspect the result, and only decisions that cannot be resolved automatically return to people. A decision recorded with its reasoning and applicability updates which future conditions still require a person and which may proceed automatically.
The entry point for an ambiguous human request needs an orchestrator that clarifies the goal, decomposes the task, selects owners, tracks progress, and integrates results.
OpenAI distinguishes manager-led orchestration, where the manager retains the conversation and calls specialists as tools, from handoffs, where a specialist becomes the active agent.
Microsoft similarly distinguishes agent-as-tools, where a primary agent keeps overall responsibility, from handoffs, where task ownership moves.
The general ability to interpret intent is not the same as permission to execute any business action directly.
If an orchestrator has the same write access as every specialist, it can bypass task-specific procedures and audit boundaries.
The safer default is for the orchestrator to define work, route it to approved executors, and integrate their progress, artifacts, validation results, and unresolved issues.
High-authority execution belongs with specialist agents or deterministic systems that have an explicit business contract.
The six steps are not necessarily a one-way sequence; evaluation results and human decisions can send work back to an earlier step.
Step 01
Set the objective and prohibited actions
Who
Accountable business owner
What
Defines the intended result, deadline, budget, permitted information, prohibited actions, decisions reserved for people, and operations that require approval because of authority.
Output
A request with success criteria, constraints, reserved decisions, and approval targets
Step 02
Break down the request and select owners
Who
Orchestrator
What
Divides the request into research, reconciliation, creation, review, or other tasks and assigns each to an owner with the required capability and authority.
Output
Task list, owners, execution order, dependencies, and required authority
Step 03
Work within the permitted scope
Who
Specialist agents, existing business systems, and human experts
What
Use only the assigned data and permissions to work toward the defined completion criteria.
Output
Results, intermediate artifacts, execution records, and reasons for stopping
Step 04
Attach evidence, completed items, and unfinished items
Who
The agent, system, or person that performed the work
What
Connects the artifact to the sources used, validation results, conditions met, unfinished items, and remaining uncertainty.
Output
An artifact with sources, validation results, completed items, and unfinished items
Step 05
Check the result against business rules
Who
Reconciliation programs, tests, and evaluation agents
What
Checks the result and execution path against criteria such as invoice totals, required contract clauses, or the correspondence between article claims and sources.
Output
Pass, fail, rerun required, insufficient evidence, or human decision required
Step 06
Own decisions reserved for people
Who
Business owner with decision rights
What
Decides priorities among objectives, acceptable risk, value conflicts, and exceptions not covered by existing rules.
Output
Decision, reasoning, applicability, exceptions, and reuse classification
Turn intent into a task contract
Record the goal, success criteria, prohibited actions, deadline, budget, required evidence, decisions reserved for people, and authority-based approvals. Keep ambiguity visible as an unresolved condition.
Select an executor and authority
Use registered capability, available data, permitted tools, and the audit method to assign work to an agent, deterministic process, or human specialist.
Return completed items, unfinished items, and evidence
Return not only the artifact, but also the conditions completed, unfinished items, reasons for stopping, uncertainty, evidence, and validation results.
Integrate and return decisions to the next run
The orchestrator groups completed, unmet, conflicting, human-reserved, and authority-approval work by goal. It preserves references to original artifacts and traces, then returns reusable conditions from human decisions to the task contract or evaluator.
Validation and Human Review
Validate artifacts by business process and show people only undecidable items and operations for which the organization retains authority
Auditing has a common record shared across tasks and acceptance criteria that differ by business process.
Common audit
Record who requested the work, which agent accepted it, what data and tools it used, which files, business records, or permissions it changed, and where it was approved, stopped, or retried. Check permissions, run counts, abnormal repetition, and irreversible actions across tasks.
Task-specific audit
Financial reconciliation, contract review, security investigation, and article production have different correctness criteria. Define expected steps, authoritative materials, tolerances, prohibited actions, and completion tests for each task.
Microsoft and Google evaluation services separately assess final task completion, instruction adherence, tool selection, inputs, use of tool outputs, and the expected execution trajectory.
This is not one universal score for a general agent.
It is a test of the contract for a particular task.
NIST's Evaluation Probes similarly begin with one bounded question: whether an AI claim is supported by a human-curated corpus of trusted documents.
“A source is cited,” “the answer matches the source,” “the source is true in the world,” and “the result is approved for business use” require separate judgments.
A practical HACW therefore needs staged evidence labels such as unverified, source-grounded, independently checked, system-validated, human-approved, contested, and expired.
Making an agent explain everything does not create effective oversight if people cannot process the volume and speed.
Approval queues and review fatigue become a new bottleneck.
Gartner's highest autonomy level has people review exceptions, audit logs, and aggregated outcomes rather than each individual decision.
Early 2026 research also describes constant inspection of AI artifacts and floods of suggestions as sources of cognitive load and hidden cost.
People do not monitor every operation. They review what changed since the last checkpoint, which items failed validation, which decisions remain human, and which operations the organization reserves for explicit approval.
A HACW should therefore retain the full run record in machine-readable form while showing people only what changed since the last checkpoint.
A decision request should contain the current goal, newly completed or failed items, significant actions, automated validation results, open issues, the consequence of taking no action, and the requested decision.
Full traces should be available for investigation, not pushed into every review.
Explicit approval for a high-risk operation is a control that reserves execution authority for the organization, not evidence that the agent is incapable of making a judgment.
Low risk
Execute automatically after validation and inspect later through sample audits.
Medium risk
Surface only exceptions, conflicting evidence, and weakly supported results in a compact decision card.
High risk
Require explicit approval immediately before publishing, sending, changing permissions, or moving money.
Undefined or contested
Stop autonomous execution and route to the accountable owner rather than filling the gap with a plausible guess.
Decision Memory
A recorded human decision becomes a rule for future work only after its applicability is validated
A decision made in Step 06 does not end as a one-time approval. Its reasoning and applicability enter the shared record, and reusable decisions return to the task contract in Step 01 or the evaluator in Step 05.
The OpenAI Agents SDK manual approval flow pauses immediately before a sensitive tool call and asks a person to approve or reject it. When a condition can be decided in code, an approval callback can also let the run continue without a human interruption.
Both mechanisms authorize an operation in a running execution. They do not necessarily turn a business decision into a rule that applies to separate future tasks.
If the reasoning and applicability do not return to a task contract, evaluator, or authorization policy, the same business conditions can call a person again next time.
Per-action approval gate
Decision unit
One tool operation in the current run
Output
Allow or reject from a person or approval program
Separate future task
The same business conditions may stop again
HACW decision update
Decision unit
A business decision that can hold across separate future tasks
Human output
Decision, reasoning, scope, exceptions, and reuse classification
Next time
Work proceeds automatically when a validated rule matches
Memory is not execution authority
Decision record
Stores the task class, preconditions, evidence, decision, reasoning, applicability, exceptions, expiry, and accountable owner. Future tasks use it as reference material.
Approved business rule
When the same decision can be tested under the same conditions, it becomes a versioned task contract, evaluator, or authorization policy. Only then can it authorize automatic execution.
Decisions retained by people
Choice of objective, value conflicts, negotiation, relationships, and acceptance of responsibility remain human decisions even when prior cases are available.
When an invoice with the same conditions arrives next month, a conventional assistant leaves a person to find the prior answer and approval, then carry them into the new task.
A HACW compares the prior reasoning, amount, counterparty, exceptions, and expiry with the new invoice. It runs automatically when an approved rule matches; otherwise it returns the prior decision to a person as reference material.
The dividing line is not agent intelligence. It is whether a decision from one completed task becomes a tested condition under which a separate future task may proceed without another human review.
Semantic similarity to a past decision is not enough to authorize a new task.
The system must compare amount, counterparty, confidentiality, applicable policy, reversibility, and expiry, then proceed only when an approved rule matches.
If the case is merely similar, the prior decision becomes reference material for a person; if information is missing, the task returns to research or reevaluation.
Reflexion and ExpeL show that agents can improve later trials by storing and retrieving reflections or extracted experience.
OpenAI Agent memory similarly carries lessons and user corrections from prior work into future runs while warning that memory can become stale and the current environment should take precedence.
AWS AgentCore Policy places the execution boundary elsewhere: conditions are evaluated outside the agent as validated, deterministic policies rather than treated as agent memory.
Taken together, these mechanisms suggest a boundary in which memory retrieves decision candidates while validated rules control automatic execution.
Three Modes
Procedure clarity, failure impact, and reversibility determine what agents execute and what people decide
HACW is not one automation method.
Even with the same work record and rule-promotion process, the tasks that run automatically and the decisions returned to people change across clear procedures, flexible work, and ideas or negotiation where human judgment remains central.
Clear procedure and judgment
Stable work such as invoice matching or routine reporting should become a deterministic business process. HACW manages authority, evidence, and exceptions in the background, while people review rule changes and unusual cases.
Flexible procedure with judgment
Contract review, incident investigation, and complex development benefit from an orchestrator and specialist agents. People own objectives, changes to evaluation criteria, and unprecedented exceptions, while reusable decisions move into task contracts or evaluators.
Ideas, negotiation, and strategy
For new ventures or negotiations, agents gather research, comparisons, options, and counterarguments. People retain direction, values, relationships, and final accountability. The audit concerns assumptions, inputs, alternatives, and decision ownership rather than one correct answer.
Risk and irreversibility must be added to this classification.
A large transfer can require approval even when every procedural step is clear, while a flexible internal brainstorm can allow substantial autonomy when the consequences are low.
The purpose of HACW is not always to remove people; it is to select the right division of responsibility for each kind of work.
The following are operating scenarios, not published deployment cases. They apply the three modes above to show how automated validation and human judgment divide differently by business process.
Month-end close
Specialists prepare entries, reconciliations, and variance analysis, then return control totals and workpapers. Normal items proceed automatically; out-of-policy variances, missing data, and approval items are escalated.
Contracts and procurement
Separate agents extract clauses, compare internal policy, and research the counterparty. The orchestrator integrates evidence and disagreement, but people choose acceptable risk and negotiation terms.
New business
Agents research markets, customers, regulation, and competition in parallel, then present options with assumptions and counterevidence. They do not declare the one correct strategy; the workspace records which values and risks people chose.
What Is Missing
Products provide orchestration, approvals, evaluation, and permissions, but no common format carries evidence and decisions across them
OpenAI, Microsoft, Google, and Anthropic each implement orchestration, specialist agents, sessions, approvals, evaluations, traces, or permission controls.
Research proposals add more of the shared-work layer: CHAP defines workspaces, participants, tasks, artifacts, and an append-only evidence log; CLEO and Pista study interfaces where people can observe, direct, work concurrently, and stop execution.
These parts have not yet converged into one mature standard or product category.
What is missing is a shared contract that carries the request objective, execution results, validation results, and human decisions across those features.
Organizations need task contracts and evaluators for each business process; a registry of agent capabilities, owners, and authority; common labels attached to artifacts and evidence; evaluation of whether the orchestrator chose the right executor; a portable audit format across products; and metrics that cap the amount of human review time.
Decision records need applicability, exceptions, expiry, and accountable ownership; promotion to a business rule needs testing, versioning, approval, and a rollback path.
Run counts are not enough to measure value.
Organizations need the total cost per validated artifact, human review time, escalation rate, decision waiting time, rework rate, defects found after automatic approval, and audit backlog, measured separately from artifact accuracy, completeness, and fitness for purpose.
Conclusion
Beyond AI agents, the question is whether organizations can narrow human review and safely reuse validated decisions
Takeaway
The comparison axis beyond AI agents is not how much work they can execute without asking.
It is whether human decisions retain their reasoning and applicability, reusable decisions become validated rules, and memory itself never becomes execution authority.
HACW becomes easier to understand when it is framed not as a place to build agents, but as the place where people and agents divide responsibility and finish work together.