Agents as Labor
A Foundational Economics and Market Design Framework for the Agent Era Why Agent Markets Will Converge on Labor Market Structures And How to Build the Infrastructure That Wins
Executive Summary
The AI agent market is experiencing a fundamental pricing miscategorization. Current commercialization models treat agents as software (priced by tokens, seats, or API calls), yet agentic systems exhibit economic characteristics that map directly to labor markets: variable performance, quality heterogeneity, task-specific matching requirements, and outcome uncertainty. This category error leaves significant value on the table and creates misaligned incentives across the value chain.
This document presents a first-principles framework demonstrating why agent markets will converge on labor market primitives. The analysis draws on Nobel Prize-winning economic theory (Akerlof on information asymmetry, Holmstrom on moral hazard, Rochet-Tirole on two-sided markets) and validates predictions against empirical market data from gig economy platforms, enterprise AI adoption surveys, and emerging outcome-based pricing models.
Core Thesis: Agents satisfy four defining characteristics that distinguish labor from software: (1) variable performance across tasks and contexts, (2) matching requirements based on capability fit rather than availability, (3) compensation that should reflect value created rather than effort expended, and (4) accountability structures with consequences for failure. Organizations that recognize this structural reality and build the corresponding infrastructure (capability taxonomies, evaluation standards, reputation systems, outcome contracts) will capture the durable value in this market.
1. The Category Error: Why Current Pricing Models Fail
1.1 Software Economics: The Default Mental Model
Software has traditionally exhibited specific economic characteristics that justify standard pricing models:
Deterministic behavior: Given identical inputs, software produces identical outputs, enabling regression testing and quality guarantees
Near-zero marginal cost: Once developed, each additional unit costs essentially nothing to produce
Feature-based differentiation: Value is captured through capability lists and version tiers
Uptime as the primary SLA: Availability, latency, and security define service quality
These characteristics support per-seat, per-usage, and tiered subscription pricing. The SaaS model that Salesforce pioneered works because the relationship between payment and value is relatively stable: you pay for access to capabilities, and the software delivers those capabilities consistently.
1.2 Agent Economics: A Structural Departure
Agentic systems break from software economics in ways that are not incremental but categorical. Analysis of agent behavior reveals four defining characteristics:
Characteristic 1: Variable Performance
Two implementations of “the same” agent can differ dramatically across tool competence, refusal patterns, hallucination rates, domain grounding, and edge case handling. This mirrors labor market dynamics: two accountants are not interchangeable despite holding identical credentials. Benchmark saturation (where multiple models score 90%+ on standard tests) obscures real-world performance variance that enterprises experience in production.
Characteristic 2: Output Uncertainty
Even with identical prompts, agents can vary due to stochastic decoding, tool latency, retrieval differences, and context drift. This is not a bug to be fixed; it is intrinsic to how these systems function. The implication is fundamental: you cannot treat agent outputs as deterministic deliverables. You must manage agents the way you manage workers, with oversight, performance measurement, and escalation paths.
Characteristic 3: Asymmetric Failure Costs
A 2% hallucination rate in marketing copy is tolerable. A 0.2% error rate in payroll, compliance, or hiring decisions can be existential. This drives the need for risk-tiered capabilities, compliance-aware routing, and contracts that incorporate downside exposure. Software SLAs built around uptime cannot capture this domain-specific risk asymmetry.
Characteristic 4: Value at Resolution
The critical distinction is that agents create value when they resolve outcomes, not when they run. A chatbot produces text. An agent changes state: files tickets, updates records, routes cases, reconciles data, triggers workflows, and escalates when needed. This is the economic signature of labor, not software.
1.3 Quantifying the Mismatch
2. Theoretical Foundations: Why Labor Economics Applies
The argument that agent markets will converge on labor market structures is not speculative. It follows directly from established economic theory that has been validated across multiple domains. Three theoretical frameworks are particularly relevant.
2.1 Asymmetric Information and the Lemons Problem
Theory: George Akerlof’s 1970 paper “The Market for ‘Lemons’“ demonstrated that in markets with information asymmetry regarding quality, buyers cannot distinguish high-quality from low-quality offerings. Sellers of high-quality goods cannot get paid fairly, so they exit, leaving only “lemons” behind. This adverse selection drives quality collapse. Akerlof received the Nobel Prize in 2001 for this work.
Application to Agent Markets: Enterprises currently cannot reliably distinguish a safe, high-performing agent from a brittle demo. Benchmark saturation (multiple models scoring 90%+ on public benchmarks) provides false confidence while masking production-relevant performance differences. The predictable outcome: buyers default to lowest-cost options or incumbent bundles, innovation commoditizes, and quality providers cannot capture appropriate value.
Required Solution: Credible quality signals. Labor markets solve this through credentials, verified work history, standardized evaluation, and warranties. Agent markets require equivalent infrastructure: capability taxonomies that specify what agents can do under what conditions, evaluation protocols that produce verifiable performance data, and reputation systems that track outcomes over time.
2.2 Principal-Agent Theory and Moral Hazard
Theory: Bengt Holmstrom’s work on moral hazard (Nobel Prize 2016) formalized how incentive contracts should be structured when effort cannot be directly observed. The key insight is the Informativeness Principle: any measure of performance that reveals information about effort should be included in compensation contracts. When effort is unobservable but outcomes are measurable, contracts should be outcome-based.
Application to Agent Markets: You cannot directly observe “effort” inside an AI model. But you can observe outcomes, tool traces, and error rates. AI agents operate autonomously with limited observability. Token-based pricing measures effort, not outcomes. The contract design implication is clear: compensation should use monitored outcome signals, including resolution success, escalation behavior, policy adherence, time to complete, and correction rates. Per-token or per-action pricing violates this principle by paying for effort proxies rather than outcomes.
Required Solution: Monitoring infrastructure and outcome contracts. This means standardized trace layers (OpenTelemetry-style observability), defined success criteria for task categories, and pricing mechanisms that tie payment to verified resolution rather than activity.
2.3 Two-Sided Markets and Platform Economics
Theory: Jean-Charles Rochet and Jean Tirole’s foundational work on two-sided markets (Journal of the European Economic Association, 2003) demonstrated that platforms serving multiple user groups face unique pricing and governance challenges. Success requires “getting both sides on board,” and the platform that solves the chicken-and-egg problem while setting pricing structure can dominate the market.
Application to Agent Markets: Agent markets are inherently two-sided (or multi-sided): supply side includes agent builders, operators, and model providers; demand side includes enterprises with tasks and constraints; the platform layer provides matching, evaluation, and contracting. Network effects compound: more agents attract more enterprises, which attract more specialized agents.
Strategic Implication: The platform that defines the capability taxonomy and evaluation standards becomes the clearinghouse. This is the current strategic window. Whoever builds the matching infrastructure captures the market-making position.
3. Empirical Validation: Market Evidence
The theoretical framework predicts specific market structures and behaviors. This section validates predictions against observed data from three sources: the gig economy (mature labor platform), enterprise AI adoption (demand-side behavior), and emerging outcome-based pricing (supply-side experimentation).
3.1 The Gig Economy as Structural Template
The gig economy provides the closest existing analog to where agent markets are heading. Both involve matching workers (human or AI) to tasks through platforms, with quality signals mediating the relationship.
Market Scale and Structure
• Global gig economy market size: $582.2B in 2025, projecting to $2.18T by 2034 (15.79% CAGR)
• Active participants: 160M+ workers globally; 38M+ in the United States alone
• Platform segment: $25.5B in 2024 for gig economy platforms specifically, growing at 20% CAGR
• Commission-based models: 70% of platform revenue, validating outcome-linked payment structures
3.2 Enterprise AI Adoption: Demand-Side Evidence
Enterprise adoption data reveals both the scale of the opportunity and the governance gaps that labor market infrastructure would address.
Adoption Velocity
85% of enterprises expected to implement AI agents by end of 2025 (multiple analyst surveys)
81% of business leaders expect AI agents deeply integrated into strategic roadmaps within 12-18 months (Microsoft Work Trend Index)
40% of enterprise applications will include task-specific agents by 2026, up from less than 5% in 2025 (Gartner)
78% of global organizations already use some form of AI tools in daily operations
Governance and Quality Concerns
75% of firms will fail at building advanced agentic architectures independently (Forrester)
85% of agentic architecture builds fail without proper governance frameworks
Over 40% of agentic AI projects may be canceled by 2027 if they lack clear value or governance (Gartner)
Trust in fully autonomous agents fell from 43% to 27% year-over-year (Capgemini)
Only 2% of firms have fully scaled AI agent deployments
Interpretation: The data shows a market moving rapidly toward universal adoption while simultaneously experiencing governance failures. This is precisely the environment where labor market infrastructure (credible quality signals, standardized evaluation, accountability structures) becomes essential for market function.
3.3 Outcome-Based Pricing: Supply-Side Experimentation
The emergence of outcome-based pricing models, particularly Sierra AI’s approach, provides early evidence that the market is shifting toward a labor economics framework.
The Sierra Model
Sierra AI, co-founded by Bret Taylor (former Salesforce co-CEO) and Clay Bavor (former Google), has pioneered outcome-based pricing for customer service agents. Key elements:
Payment trigger: Companies pay only when a customer issue is resolved through the AI agent
Cost benchmark: A typical customer service call costs $10-20, mainly labor. Sierra charges a portion of these savings.
Performance results: 50-90% of customer service interactions fully automated; satisfaction scores up to 4.6-4.7 out of 5
Incentive alignment: Sierra only profits when technology delivers value, creating motivation for continuous improvement
Economic Logic
Taylor’s articulation captures the thesis precisely: “We think outcome-based pricing is the future of software. With AI, we finally have technology that isn’t just making us more productive but actually doing the job. It’s actually finishing the job.” This is the economic signature of labor: payment for completed work, not for access or activity.
4. What a Mature Agent Labor Market Requires
The transition from software economics to labor economics requires specific infrastructure that does not yet exist at scale. This section specifies the primitives without which the market will not stabilize at high quality.
4.1 Capability-Based Matching
Instead of “buy our agent,” the buyer should be able to specify: “Match me to the right agent for this task type, in this context, under these constraints, with this risk tolerance and this SLA.” This requires a capability taxonomy that is:
Structured: Machine-readable specifications, not marketing descriptions
Compositional: Capabilities compose into roles and workflows
Contextual: Capability depends on tool access, data scope, and jurisdiction
Measurable: Each capability has evaluation protocols with verifiable metrics
Capability Schema Structure
A capability is not “can do payroll.” It is a structured specification including:
Task class: e.g., “classify inbound HR case,” “reconcile time-off balances,” “draft compliant offer letter.”
Inputs: Data types, schema, language requirements
Tools: HRIS actions allowed, ticketing systems, email, knowledge base retrieval
Constraints: Policy compliance, privacy requirements, risk tier
Success criteria: What constitutes “done,” what is “acceptable.”
Uncertainty handling: When to escalate, when to ask clarifying questions
Performance envelope: Latency requirements, throughput constraints
4.2 Quality Signals and Reputation
Labor markets run on signals: credentials, references, work history, verified outcomes. Agent markets currently have weak signals (benchmarks do not equal enterprise reality), creating the lemons dynamics Akerlof identified.
A credible agent reputation system must include:
Work Ledger: A signed record of tasks attempted, outcomes achieved, and escalations made (privacy-preserving, enterprise-controlled)
Context Tags: Outcomes comparable only within similar contexts (tools, data, domain, risk level)
Verification: Outcomes objectively verified (tests, audits, human sign-off, downstream system confirmation)
Adversarial Robustness: Prevention of metric gaming (Goodhart effects)
Portability: Reputation not locked to one vendor suite
4.3 Outcome-Based Pricing
Outcome pricing becomes possible when you can: define “done,” verify it, attribute success or failure clearly, and manage risk. Examples by domain:
Customer support: $X per resolved case, with “resolved” defined as “no reopen within N days plus CSAT threshold”
Recruiting: $X per qualified slate delivered, defined as “meets rubric plus hiring manager acceptance plus adverse impact checks completed”
Finance: $X per reconciled account, defined by “no exceptions in audit plus downstream close step passes”
Risk-adjusted pricing applies: high-risk tasks (regulatory exposure) require higher prices and stricter SLAs; low-risk tasks can be cheaper and more automated.
4.4 Performance Accountability and SLAs
Current SLAs in AI typically focus on uptime, latency, and availability. Labor-market SLAs are quality, error rates, rework rates, time-to-resolution, compliance adherence, and escalation behavior.
When an agent fails, responsibility is often fuzzy: model provider? Agent builder? Integrator? Enterprise user? A mature market requires explicit accountability allocation, like subcontracting in services. Contracts should define:
Who owns the workflow definition
Who owns monitoring
Who owns safety policies
Who pays for remediation
5. The Agent Labor Stack: Infrastructure Layers
Think of this as the “AWS of labor markets,” not the “app store of bots.” The stack has six layers, with value and margin concentrating at the middle layers as the market matures.
Key Insight: Most current products sit at Layers 0-1 or bundle Layers 1-2 inside suites. The power and margin concentrate at Layers 2-5 once the market matures. This is where knowledge graphs and ontologies become decisive: they are the substrate for capability registries, evaluation frameworks, and performance tracking.
6. Strategic Implications: Positioning for Market Design
For organizations with existing ontology or knowledge graph capabilities for skills, roles, and workflows, the strategic opportunity is substantial. These organizations are not “adjacent” to this market; they are holding the substrate that makes it function.
6.1 The Knowledge Graph Advantage
A mature capability taxonomy is essentially an ontology. Organizations with existing structured representations of work (tasks, skills, roles, competencies, outcomes) have four strategic plays:
1. Become the capability standard. Publish and implement a capability schema for your domain: tasks linked to skills linked to roles linked to tools linked to constraints linked to success metrics. Make it compositional and machine-readable.
2. Own evaluation. Build domain-specific “SWE-bench equivalents” for your vertical: curated task suites, simulation environments, verification harnesses, red-team cases. Whoever defines the evaluation standards owns the quality narrative.
3. Build the reputation ledger. For each agent instance: track verified outcomes, tag context, compute performance distributions, publish “credentials” internally or to a market.
4. Move pricing up the stack. Offer outcome-backed pricing with warranties: “We guarantee X% resolution; we pay penalties if not met.” The winners will be those who can underwrite outcomes because they can measure and control them.
6.2 The Market-Making Position
Once you have capability + evaluation + reputation, matching becomes natural: route tasks to the best agent for that context, diversify supply, build switching costs around trust rather than integration. This is how you avoid becoming a “wrapper” and instead become infrastructure.
The parallel to gig economy platforms is instructive. Platforms like Uber, Upwork, and DoorDash capture 5-20% take rates not because they own the workers or the customers, but because they provide the matching infrastructure and trust layer. The equivalent position in agent markets goes to whoever provides:
• The capability language that buyers and sellers use to specify requirements
• The evaluation benchmarks that define quality
• The reputation data that enables trust
• The contract templates that allocate risk
6.3 Window of Opportunity
The market is being designed now. With 85% of enterprises adopting agents by end of 2025, and 40% of applications including task-specific agents by 2026, the infrastructure layer will solidify within 12-18 months. Organizations that define the primitives (capabilities, credentials, evaluation, contracts) during this window become the standards-setters. Standards-setters accrue the durable power and margin in platform markets.
7. Risks, Failure Modes, and Design Responses
7.1 Goodhart’s Law and Metric Gaming
Risk: If you pay per resolved case, agents may “close” prematurely to maximize the metric.
Design Response: Use reopen windows (resolution not counted until N days without reopen). Weight customer satisfaction and downstream correctness. Penalize rework and policy violations. Design metrics that are robust to simple optimization strategies.
7.2 Hidden Labor
Risk: Vendors may quietly route hard cases to humans while claiming agent performance.
Design Response: Require provenance of decisions and actions. Instrument escalation paths. Separate “assisted resolution” from “autonomous resolution” in metrics and pricing.
7.3 Security and Tool Misuse
Risk: Agents are attack surfaces with access to sensitive systems.
Design Response: Principle of least privilege for tool access. Action signing and approval thresholds for high-risk acts. Trace-based anomaly detection.
7.4 Liability Ambiguity
Risk: Without clear accountability allocation, adoption stalls or litigation increases.
Design Response: Contractual responsibility allocation and insurance-like structures. Auditable logs and traceability aligned with emerging governance expectations (NIST AI RMF, EU AI Act frameworks).
8. Conclusion: The Convergence Thesis
The argument presented in this document rests on a structural claim: agents exhibit the economic characteristics of labor, not software, and markets eventually price assets according to their actual characteristics. The theoretical foundations (Akerlof, Holmstrom, Rochet-Tirole) predict that information asymmetry requires quality signals, unobservable effort requires outcome contracts, and two-sided markets require platform infrastructure.
The empirical evidence validates this trajectory. The gig economy demonstrates that digital labor markets at scale require matching infrastructure, reputation systems, and outcome-based payment. Enterprise AI adoption shows both rapid scaling and governance failures that labor market primitives would address. Early movers like Sierra demonstrate that outcome-based pricing is viable when measurement infrastructure exists.
The strategic implication is clear: the value in agent markets will accrue not to model providers (commoditizing) or to simple wrappers (undifferentiated), but to the infrastructure layer that makes the market function. This infrastructure is built on structured representations of work: capability taxonomies, evaluation frameworks, reputation systems, and outcome contracts.
Organizations with existing knowledge graphs and ontologies for skills, roles, tasks, and outcomes are positioned to capture this opportunity. They hold the substrate that agent labor markets require. The strategic window is 12-18 months. The organizations that define the primitives during this period become the standards-setters, and standards-setters in platform markets accrue the durable power and margin.
The bottom line: Agents are labor, not software. Markets will price them accordingly. The infrastructure that enables labor market function (matching, reputation, outcome contracts, accountability) represents the strategic high ground. Build it now.




