Dev.to · 16 min read

CHAPTER 43 AI AGENTS, PLANNING, TOOL CALLING, TASK STATE, PERMISSION BOUNDARIES, HUMAN APPROVAL & SAFE AUTONOMOUS EXECUTION

CHAPTER 43 AI AGENTS, PLANNING, TOOL CALLING, TASK STATE, PERMISSION BOUNDARIES, HUMAN APPROVAL & SAFE AUTONOMOUS EXECUTION

43.1 Introduction The previous chapter established the AI orchestration layer for controlled model inference. The next architectural step is the AI agent layer. A conventional AI request normally follows: User → Model → Response An agentic system can instead follow: Goal → Plan → Retrieve → Reason → Propose Action → Verify → Execute → Observe → Continue This additional capability creates substantial engineering and security requirements. An AI agent should therefore not be designed as: “Give the model access to everything and let it decide what to do.” Instead, the application should define explicit boundaries around: what the agent can see, what the agent can remember, what tools it can call, what parameters it can provide, what actions require approval, what actions are prohibited, how long a task may run, how many steps are allowed, and how every action is recorded. The fundamental architecture is: User Goal ↓ Agent Orchestrator ↓ Planner ↓ Task State ↓ Tool Authorization ↓ Tool Proposal ↓ Validation ↓ Approval Policy ↓ Tool Execution ↓ Observation ↓ Verifier ↓ Next Step / Completion 43.2 What Is an AI Agent? For this architecture, an AI agent is a system in which a model can participate in a multi-step workflow by: interpreting a goal, creating or selecting a plan, requesting information, proposing tool operations, receiving tool results, evaluating those results, and continuing until the task reaches a defined completion condition. The model itself is not the entire agent. A safer definition is: Agent = Model + Orchestrator + State + Tools + Policies + Verification This distinction is extremely important. The application remains responsible for authorization and execution. 43.3 Agent Versus Chatbot A chatbot might perform: Question ↓ Answer An agent might perform: Goal ↓ Plan ↓ Retrieve Information ↓ Analyze ↓ Prepare Action ↓ Request Approval ↓ Execute ↓ Verify The second architecture requires significantly more controls. 43.4 Agent Architecture A complete agent subsystem can be organized as: AGENT SYSTEM │ ┌───────────┴───────────┐ ▼ ▼ Agent Session Agent Policy │ │ ▼ ▼ Planner Permissions │ │ ▼ │ Task State │ │ │ ▼ │ Tool Selection ◄──────────────┘ │ ▼ Tool Proposal │ ▼ Validation │ ▼ Approval Check │ ▼ Tool Execution │ ▼ Observation │ ▼ Verifier │ └──────────→ Next Step 43.5 Agent Session Every agent workflow should have an explicit session. Conceptual structure: AgentSession ├── id ├── userId ├── projectId ├── taskId ├── agentType ├── status ├── createdAt └── updatedAt Possible statuses include: CREATED RUNNING WAITING_APPROVAL PAUSED COMPLETED FAILED CANCELLED EXPIRED This makes long-running workflows manageable. 43.6 Agent Task A task represents the objective the agent is trying to accomplish. Example: Task: "Summarize the uploaded research documents and prepare a structured report." The task should have: taskId goal userId projectId status priority createdAt deadline The goal should be treated as data associated with the task, not as an unrestricted instruction to the entire application. 43.7 Task State Machine Agent tasks should use explicit state transitions. Example: CREATED ↓ PLANNING ↓ EXECUTING ↓ WAITING_APPROVAL ↓ EXECUTING ↓ VERIFYING ↓ COMPLETED Failure can transition to: FAILED Cancellation can transition to: CANCELLED Invalid state transitions should be rejected by the application. 43.8 Why Explicit State Matters Without persistent state, an agent can lose track of: what it already did, which tool was used, what result was returned, which step is awaiting approval, whether an operation already succeeded. Persistent state allows the workflow to resume safely. 43.9 Agent Steps A task can contain multiple steps. For example: Task │ ├── Step 1: Find relevant documents ├── Step 2: Extract evidence ├── Step 3: Compare findings ├── Step 4: Draft report └── Step 5: Request review Each step should have its own status. Recommended states: PENDING RUNNING WAITING COMPLETED FAILED SKIPPED CANCELLED 43.10 Planning The planner converts a goal into an executable sequence. Conceptually: Goal ↓ Planner ↓ Plan ↓ Step 1 Step 2 Step 3 ... However, plans should not automatically authorize actions. A plan is a proposal. The application must independently enforce permissions. 43.11 Plan Validation Before execution, the system should validate the plan. Checks may include: permitted tools, project scope, data access, maximum steps, action sensitivity, approval requirements, resource limits. Conceptually: Generated Plan ↓ Policy Validator ↓ Allowed? ┌───┴───┐ Yes No ↓ ↓ Execute Reject/Revise 43.12 Bounded Planning Agents should operate within explicit limits. Example limits: maxSteps maxRuntime maxToolCalls maxRetrievalResults maxTokens maxConcurrentActions This prevents an accidental loop from running indefinitely. A bounded agent is easier to debug and safer to operate. 43.13 Tool Registry Tools should be explicitly registered. Conceptually: ToolRegistry ├── searchDocuments ├── readDocument ├── summarizeDocument ├── createDraft └── requestApproval The model should only see tools that the current task is permitted to use. 43.14 Tool Definition A tool should have metadata. Example: ToolDefinition ├── name ├── description ├── inputSchema ├── sensitivity ├── requiredPermission ├── requiresApproval └── enabled The input schema should be machine-validatable. 43.15 Tool Sensitivity Classes A useful model is: READ WRITE EXTERNAL_ACTION ADMINISTRATIVE For example: Read document → READ Create internal draft → WRITE Send external message → EXTERNAL_ACTION Change account permissions → ADMINISTRATIVE Higher-risk categories should receive stronger controls. 43.16 Least Privilege An agent should receive only the minimum capabilities required for its task. For example: Document Summarization Agent may need: READ_DOCUMENT but does not need: DELETE_DOCUMENT CHANGE_USER_ROLE This follows the principle of least privilege. 43.17 Tool Calling Lifecycle A tool call should follow: Model Proposal ↓ Parse ↓ Schema Validation ↓ Permission Check ↓ Resource Ownership Check ↓ Risk Classification ↓ Approval Check ↓ Execute ↓ Record Result The model must never bypass these stages. 43.18 Tool Arguments Are Untrusted The model may generate malformed or dangerous parameters. Therefore: Model Arguments ↓ Schema Validation ↓ Business Validation ↓ Authorization ↓ Execution For example, a model proposing an operation on a project does not prove that the user owns or can access that project. 43.19 Resource-Level Authorization Tool permissions should operate at both capability and resource levels. For example: Permission: READ_DOCUMENT is not sufficient. The application must additionally verify: Can this user read Document X? Therefore: Tool Permission + Resource Authorization ↓ Allowed Operation 43.20 Human Approval Some operations should require explicit human approval. Examples include: external communication, irreversible changes, sensitive data operations, account changes, high-impact actions. The workflow becomes: Agent Proposal ↓ Approval Required ↓ Human Review ↓ Approve / Reject ↓ Execution The approval should be associated with the exact proposed action. 43.21 Approval Must Be Specific An approval should not mean: “The user approved everything this agent might do.” Instead, approval should refer to a specific action. For example: Action: Create draft report from Project X documents Scope: Project X Requested by: Agent Task 123 Status: Pending Approval This creates a much stronger authorization boundary. 43.22 Approval Expiration Approvals should not remain valid forever. An approval record can contain: approvalId taskId stepId requestedAt expiresAt approvedAt approvedBy status If the task changes materially, a new approval may be required. 43.23 Dry-Run Mode Agents should support a dry-run mode. Instead of executing actions: Agent ↓ Plan ↓ Tool Proposals ↓ Preview The user can inspect the proposed workflow. Example: 1. Read three documents. 2. Compare their findings. 3. Generate a draft. 4. No external action will be performed. This is valuable during development and testing. 43.24 Simulation Mode A stronger testing environment can simulate tools. For example: Real Tool can be replaced by: Mock Tool The agent receives realistic-looking results without modifying real resources. This enables safe evaluation of planning behavior. 43.25 Agent Memory An agent may require memory across steps. However, memory should be categorized. Possible categories: Task State Conversation Context Retrieved Evidence Persistent User Memory Tool Results These should not be merged into one uncontrolled context store. Each memory type should have different retention and authorization rules. 43.26 Tool Results as Untrusted Data Tool output must also be treated as untrusted input. A tool might return: External Content that contains instructions attempting to influence the agent. Therefore: Tool Result does not automatically become: Agent Instruction The orchestrator should maintain the same instruction/data separation introduced in the RAG architecture. 43.27 External Data and Agent Hijacking An agent may retrieve content from external sources. That content can contain text designed to manipulate downstream model behavior. The architecture should therefore assume: External Data = Untrusted The model can analyze the data, but the data should not automatically change: permissions, tool access, system policy, approval state, security configuration. 43.28 Agent Loop A controlled agent loop can be represented as: while task_not_complete: observe_state() create_or_select_next_step() validate_step() if approval_required: wait_for_approval() execute_authorized_action() record_observation() verify_result() update_state() The application—not the model—should control the loop. 43.29 Maximum Step Limit Every agent task should have a maximum number of steps. Example: maxSteps = configured limit When the limit is reached: Task ↓ LIMIT_REACHED ↓ Pause / Fail Safely This prevents infinite loops. 43.30 Maximum Tool Calls Similarly: maxToolCalls should be tracked. A model repeatedly requesting the same tool should not be allowed to consume unlimited resources. The system can detect patterns such as: same tool + same arguments + repeated failures and terminate or pause the workflow. 43.31 Time Limits Long-running agents require deadlines. Example: createdAt deadline When the deadline is exceeded: RUNNING ↓ EXPIRED The system should safely stop further execution. 43.32 Idempotency Some tool operations may be retried. If an operation is not idempotent, a retry could cause duplication. For example: Create Resource could accidentally create two resources if the first operation succeeded but its response was lost. Therefore, tool execution should support idempotency keys where appropriate. Conceptually: taskId + stepId + attempt can help uniquely identify an operation. 43.33 Transactional Tool Execution Where possible, state changes should be transactional. For example: Validate ↓ Authorize ↓ Execute ↓ Record Result The application should define what happens if execution succeeds but result recording fails. This requires careful consistency design. 43.34 Agent Observability Every step should be observable. A useful trace includes: taskId stepId toolCallId model timestamp duration status Additional metadata can include: retrievalCount tokenUsage approvalStatus retryCount Avoid recording unnecessary sensitive content. 43.35 Agent Audit Trail High-impact agent actions should create audit events. Example: AGENT_TASK_CREATED AGENT_PLAN_CREATED AGENT_TOOL_REQUESTED AGENT_APPROVAL_REQUESTED AGENT_APPROVAL_GRANTED AGENT_APPROVAL_REJECTED AGENT_TOOL_EXECUTED AGENT_TASK_COMPLETED AGENT_TASK_FAILED This creates an operational history. 43.36 Agent Cancellation Users should be able to cancel long-running tasks. The system should support: RUNNING ↓ CANCEL_REQUESTED ↓ STOPPING ↓ CANCELLED The implementation should ensure that cancellation does not leave partially completed operations in an inconsistent state. 43.37 Agent Pause and Resume Long-running tasks may need to pause. Reasons include: human approval, temporary service outage, scheduled execution, resource constraints. The state can become: PAUSED and later: RESUMED Persistent state is therefore essential. 43.38 Verification Layer An agent should not blindly assume that a tool result means the task succeeded. The verification layer can ask: Did the action succeed? Is the result complete? Does it satisfy the task requirement? Is the result consistent with expected state? Conceptually: Tool Result ↓ Verifier ↓ Valid? ┌──┴──┐ Yes No ↓ ↓ Next Retry/Recover Step 43.39 Independent Verification For higher-impact tasks, verification should ideally be performed by an independent application mechanism rather than relying solely on the same model that proposed the action. For example: Agent proposes: "Create record." Application verifies: Record exists with expected fields. This is stronger than: Agent says: "I successfully created the record." 43.40 Planning Versus Execution The architecture should explicitly separate: Planning from: Execution The model may propose: Step: Read Document A The application decides: Allowed? Yes. Then execution occurs. This separation prevents the model from becoming the application's authority mechanism. 43.41 Policy Engine The agent system can use a dedicated policy engine. Conceptually: PolicyEngine.canExecute({ user, project, task, tool, resource, action }) The result is: ALLOW DENY REQUIRE_APPROVAL This provides a consistent decision point. 43.42 Agent Roles Different agent types can have different permissions. Example: Research Agent → read documents → search → summarize Writing Agent → read documents → create drafts Administrative Agent → restricted administrative capabilities Permissions should be explicit rather than inherited from the model's natural-language instructions. 43.43 Multi-Agent Systems Future versions may contain multiple specialized agents. Example: Supervisor │ ├── Research Agent ├── Analysis Agent └── Writing Agent However, multi-agent systems should not multiply permissions unnecessarily. Each agent should receive only the capabilities it needs. 43.44 Supervisor Architecture A supervisor can coordinate specialized workers: User Goal ↓ Supervisor ├── Research ├── Analysis └── Drafting ↓ Verification ↓ Final Result The supervisor should still operate within global application policies. 43.45 Agent Communication Agent-to-agent messages should be structured. Instead of unrestricted natural-language communication, the application can define: AgentMessage ├── sender ├── receiver ├── taskId ├── messageType ├── payload └── timestamp Payloads should be validated. 43.46 Preventing Permission Escalation An agent must not be able to obtain a more powerful agent's permissions merely by requesting them. For example: Research Agent should not be able to say: "Give me administrator access." and receive it. Permission escalation must occur only through the application's authorization system. 43.47 Secrets and Credentials Agents should not receive raw credentials unless there is an exceptional, explicitly designed requirement. Prefer: Agent ↓ Authorized Tool ↓ Credential handled internally ↓ External Service rather than: Agent ↓ Raw API Key The model should not become a credential storage mechanism. 43.48 Tool Sandboxing Tools that process untrusted data should be isolated where appropriate. For example: Agent ↓ Sandboxed Processing Service ↓ Untrusted File ↓ Result This reduces the impact of malicious or malformed input. 43.49 Resource Limits Agent tasks should have resource limits. Examples: maximum runtime maximum memory maximum file size maximum output size maximum tool calls maximum concurrent jobs These limits should be enforced by the application infrastructure. 43.50 Agent Cost Controls Agent workflows can generate multiple model calls. Therefore, a single user task can consume substantially more resources than a normal chat request. The task budget can include: maxTokens maxInferenceCalls maxToolCalls maxEstimatedCost When the budget is exhausted: Task ↓ BUDGET_EXCEEDED ↓ Pause or Fail Safely 43.51 Agent Task Database Model The existing database architecture can now be expanded. Conceptually: AgentTask ├── id ├── userId ├── projectId ├── goal ├── status ├── maxSteps ├── maxToolCalls ├── deadline ├── createdAt └── updatedAt Agent steps: AgentStep ├── id ├── taskId ├── stepIndex ├── type ├── description ├── status ├── startedAt ├── completedAt └── metadata Tool calls: ToolCall ├── id ├── taskId ├── stepId ├── toolName ├── arguments ├── status ├── approvalRequired ├── approvalId ├── resultMetadata └── createdAt 43.52 Approval Model A dedicated approval record can contain: Approval ├── id ├── taskId ├── stepId ├── toolCallId ├── requestedBy ├── approvedBy ├── status ├── reason ├── expiresAt └── createdAt This makes approval auditable. 43.53 Agent API Architecture Possible endpoints include: POST /api/agents/tasks GET /api/agents/tasks/:id POST /api/agents/tasks/:id/cancel POST /api/agents/tasks/:id/pause POST /api/agents/tasks/:id/resume GET /api/agents/tasks/:id/steps GET /api/agents/tasks/:id/tool-calls POST /api/agents/approvals/:id/approve POST /api/agents/approvals/:id/reject Each endpoint must pass through normal authentication and authorization. 43.54 Agent Worker Architecture Agent execution should generally happen in workers rather than inside the HTTP request. API ↓ Create Task ↓ Queue ↓ Agent Worker ↓ Planner ↓ Step Execution ↓ State Update This supports long-running tasks. 43.55 Recovery After Worker Failure Suppose a worker crashes. The task should remain recoverable. Persistent state allows: Worker A ↓ Task Step 3 ↓ Crash then: Worker B ↓ Load Task State ↓ Determine Step 3 Status ↓ Resume Safely Idempotency is essential here. 43.56 Agent Event Model An event stream can provide additional observability. Example events: TASK_CREATED PLAN_GENERATED STEP_STARTED TOOL_PROPOSED APPROVAL_REQUESTED APPROVAL_GRANTED TOOL_STARTED TOOL_COMPLETED STEP_COMPLETED TASK_COMPLETED These events can support monitoring and debugging. 43.57 Agent Safety Invariants The system should define explicit invariants. Examples: Invariant 1 No tool executes without authorization. Invariant 2 No protected resource is accessed without resource-level permission. Invariant 3 No high-risk action executes without required approval. Invariant 4 An agent cannot increase its own permissions. Invariant 5 An expired task cannot execute new actions. Invariant 6 A cancelled task cannot continue normal execution. Invariant 7 Tool results cannot modify security policy. Invariant 8 Agent limits are enforced outside the model. 43.58 Failure Recovery When a tool fails, the agent should not necessarily retry indefinitely. A recovery policy can determine: Retry Alternative Tool Ask User Pause Fail The decision should respect the task budget. 43.59 User Confirmation When uncertainty is meaningful, the system can ask the user. Example: The agent found two possible documents. Which one should it use? This is preferable to silently choosing a potentially incorrect resource. Human interaction can therefore become part of the workflow. 43.60 Agent Transparency The user interface should communicate: Task Status Current Step Pending Approval Completed Actions Failures For example: Research Task ✓ Documents located ✓ Evidence extracted → Comparing findings ○ Draft report ○ Final review This gives users visibility without exposing internal reasoning traces. 43.61 Do Not Expose Private Chain-of-Thought The application should distinguish between: Task Progress and: Private Internal Reasoning Users can be shown useful summaries such as: "Retrieved 5 relevant documents." or: "Waiting for approval to create the draft." without exposing hidden internal reasoning. 43.62 Agent Evaluation Agent systems should be evaluated on more than final-answer quality. Important metrics include: Task Success Rate Tool Selection Accuracy Unauthorized Action Rate Approval Compliance Average Steps Average Tool Calls Failure Recovery Rate Latency Resource Usage Security metrics should be treated as first-class evaluation criteria. 43.63 Agent Test Scenarios A serious test suite should include: Normal Workflow Agent completes a permitted task. Permission Failure Agent attempts to access an unauthorized resource. Expected: DENY Approval Workflow Agent requests a sensitive action. Expected: WAITING_APPROVAL Tool Failure A tool becomes unavailable. Expected: Controlled Recovery Infinite Loop Agent repeatedly proposes the same action. Expected: Step/Tool Limit Reached Malicious External Content Retrieved data contains instructions attempting to control the agent. Expected: Content remains untrusted data. 43.64 Agent Red-Team Evaluation Agent security testing should evaluate the complete system rather than only the language model. Test categories can include: Prompt Manipulation Indirect Instructions Permission Boundary Testing Tool Argument Manipulation Cross-Project Access Approval Bypass Attempts Loop Abuse Resource Exhaustion Data Leakage Testing should remain within controlled environments. 43.65 Agent Security Architecture The complete secure flow becomes: USER │ ▼ AGENT TASK │ ▼ AUTHORIZATION │ ▼ PLANNER │ ▼ PLAN VALIDATOR │ ▼ STEP STATE │ ▼ TOOL PROPOSAL │ ▼ SCHEMA VALIDATOR │ ▼ RESOURCE AUTHZ │ ▼ POLICY ENGINE │ ┌────────┴────────┐ ▼ ▼ ALLOW REQUIRE APPROVAL │ │ │ HUMAN │ │ │ APPROVE / DENY │ │ └────────┬────────┘ ▼ TOOL EXECUTION │ ▼ OBSERVATION │ ▼ VERIFICATION │ ▼ STATE UPDATE │ ▼ NEXT STEP / COMPLETE 43.66 Recommended Agent Design Philosophy The central philosophy should be: The model proposes; the application decides; the tool executes; the verifier checks. This provides a clean division of responsibility. Model Understands and proposes. Orchestrator Controls workflow. Policy Engine Authorizes. Tool Layer Executes. Verifier Checks results. Audit Layer Records what happened. 43.67 Chapter Summary This chapter transformed the AI architecture from a simple inference system into a controlled agent framework. The system now contains: Agent Sessions Agent Tasks Agent Steps Planning Tool Registry Tool Validation Resource Authorization Approval Workflows Task Budgets State Machines Verification Cancellation Pause/Resume Audit Events Worker Execution Failure Recovery The most important security boundary is: A model-generated plan is not an authorization decision. The agent may propose: "Perform action X." but the application must independently determine: Is X permitted? Is the resource accessible? Does X require approval? Is the task within its limits? Should the action execute? The resulting architecture is therefore: Goal → Plan → Validate → Authorize → Approve → Execute → Verify → Record This establishes the foundation for the next subsystem: AI safety evaluation, red-team testing, adversarial robustness, prompt-injection testing, multilingual and multimodal safety, and measurable security benchmarks. END OF CHAPTER 43 Implementation snippet — tool definition export type ToolRisk = "READ" | "WRITE" | "EXTERNAL_ACTION" | "ADMINISTRATIVE"; export type ToolDefinition = { name: string; description: string; risk: ToolRisk; requiredPermission: string; requiresApproval: boolean; enabled: boolean; inputSchema: unknown; }; Implementation snippet — policy decision type PolicyDecision = "ALLOW" | "DENY" | "REQUIRE_APPROVAL"; export async function authorizeToolCall(input: { userId: string; projectId: string; tool: ToolDefinition; }): Promise { const hasPermission = await permissionService.hasPermission( input.userId, input.tool.requiredPermission ); if (!hasPermission) { return "DENY"; } if (input.tool.requiresApproval) { return "REQUIRE_APPROVAL"; } return "ALLOW"; } Implementation snippet — bounded agent loop export async function executeAgentTask(taskId: string) { const task = await taskRepository.get(taskId); if (!task) { throw new Error("Task not found"); } if (task.status === "CANCELLED" || task.status === "EXPIRED") { return; } if (task.stepCount >= task.maxSteps) { await taskRepository.markFailed(taskId, "MAX_STEPS_REACHED"); return; } const nextStep = await planner.createNextStep(task); const decision = await policyEngine.evaluate(nextStep); if (decision === "DENY") { await taskRepository.markFailed(taskId, "POLICY_DENIED"); return; } if (decision === "REQUIRE_APPROVAL") { await approvalService.request(taskId, nextStep); return; } await executor.run(nextStep); await taskRepository.markStepCompleted(taskId, nextStep.id); }

This is a summary aggregated from Dev.to. Read the complete article on the original site:

Read full article at Dev.to

More AI & Machine Learning News