top of page

Humanoid Deployment Evidence Report 2026: What Enterprise Pilots Actually Delivered (2024–2026)

A Buyer Intelligence Review of Verified Enterprise Deployments, Operational Outcomes and Procurement Readiness

Physical AI Journal Research Team  —  Independent analysis by Sekason Research Limited, United Kingdom


Executive Summary

Automation and procurement buyers evaluating humanoid robot investment in 2026 must answer a single question: which platforms have produced sufficient evidence to justify capital commitment, and under what operational conditions? Global production exceeded 20,000 units in 2025 — a tenfold increase from fewer than 2,000 units in 2024 — yet only approximately 10% entered real-world operational environments (Interact Analysis, July 2026), and of those, only two deployments have produced independently corroborated results capable of supporting a procurement business case.


This report classifies every publicly reported humanoid deployment by evidence quality, identifies where ROI is currently defensible, and provides the procurement framework to determine whether to deploy now, observe, or wait.


Two deployments have produced independently corroborated operational results as of mid-2026. Figure AI's deployment of its Figure 02 platform at BMW Group Plant Spartanburg in South Carolina delivered 30,000+ BMW X3 vehicles supported, 90,000+ sheet-metal parts handled, and 1,250+ operational hours over an 11-month programme — confirmed by a BMW Group press release, making it the first independently verified commercial humanoid production-line result in automotive manufacturing history. Agility Robotics' Digit moved more than 100,000 totes at GXO Logistics' facility in Flowery Branch, Georgia, across a live commercial deployment, a figure company-announced and corroborated by multiple trade press sources, though no independent cost-per-cycle audit has been published.


Every other enterprise humanoid deployment publicly reported as of mid-2026 — including Apptronik Apollo at Mercedes-Benz, Boston Dynamics Atlas at Hyundai's Robotics Metaplant Application Centre, and Tesla Optimus at Fremont — remains at the announcement, confirmed pilot, or internal training stage, with no independently published operational KPIs.


Tesla's own CEO confirmed on the Q4 2025 earnings call that Optimus units were

"primarily for learning, not productive tasks."

Procurement buyers should apply a single decision gate before any capital commitment: verified deployment evidence in the same task category, from the same vendor, at a comparable scale. Where that evidence exists — structured, repetitive material handling in controlled automotive or logistics environments — a pilot investment is defensible. Where it does not, the appropriate position is structured market monitoring.


The Deployment Evidence Score, introduced in Section 3, provides a five-level classification framework for every known deployment, enabling procurement teams to distinguish independently verified commercial results from announcements, vendor-reported claims, and internal testing programmes.


Humanoid robot on factory conveyor beside crate, with report cover text: Humanoid Deployment Evidence Report 2026.

1. Executive Intelligence Synthesis

Only two enterprise humanoid deployments have produced independently corroborated operational KPIs as of mid-2026 — a finding that defines the scope of every defensible procurement business case in this sector. Five signals derived from verified research translate this finding into specific guidance.


Signal 1: The industry is manufacturing faster than it is deploying. 

Of the 20,000+ humanoid robots produced globally in 2025, approximately 10% entered real-world operational environments (Interact Analysis, July 2026). Production volume, funding announcements, and market forecasts are not deployment proxies. Procurement teams should treat them as three separate data sets.


Signal 2: Only two deployments have produced independently corroborated operational KPIs. 

The BMW / Figure AI programme (Level 5 — Independently Verified Commercial Success) and the GXO / Agility Robotics programme (Level 4 — Measurable Outcomes Published) are the only cases in the public record where a customer or corroborating source has confirmed measurable throughput results. All other deployments reviewed in this report are at Level 2 or below.


Signal 3: The evidence base is narrow and task-specific.

Both validated deployments involve structured, repetitive material handling in controlled factory or logistics environments. Neither provides evidence for general-purpose operation, diverse task execution, or performance in unstructured environments. Vendor claims about broader applicability are not supported by the current evidence record.


Signal 4: Cost structure is systematically underrepresented in market coverage. 

Hardware purchase price represents 55–65% of the true five-year cost of a humanoid deployment (Association for Advancing Automation, 2025 benchmarking data). Integration, maintenance, training, safety infrastructure, and supervision costs account for the remainder. Business cases built on hardware price alone will be materially incorrect.


Signal 5: Regulatory infrastructure has not caught up with commercial deployment. 

ISO 25785-1 — the dedicated safety standard for dynamically stable humanoid robots — remains at Working Draft stage as of mid-2026, with no confirmed publication date. Buyers cannot obtain third-party certification of a humanoid deployment under a published humanoid-specific safety standard.


Buy / Observe / Wait Matrix

  • Deploy Now — Structured logistics tote-handling; automotive body-shop sheet-metal loading at fixed workstations. Conditions that must all be met: high utilisation (double shift or better), vendor with at least Level 4 evidence in the same task category, integration partner with prior humanoid experience, and total cost modelled correctly against the A3 benchmark.

  • Observe — Complex assembly, warehouse picking of varied SKUs, flexible manipulation workflows, any environment where the closest comparable deployment is Level 2 or below.

  • Wait — Healthcare, construction, hospitality, general-purpose retail, any application requiring unsupervised autonomous operation in unstructured environments. No commercially verified deployment evidence exists in these categories.


2. Platform & Market Landscape


How mature is the humanoid robotics market today?

The market is in early commercial transition. Global production exceeded 20,000 units in 2025, but only approximately 10% entered real-world operational environments (Interact Analysis, July 2026). Verified enterprise deployments with published operational outcomes remain confined to two programmes. Buyers should treat production figures, funding announcements, and analyst forecasts as three distinct data sets — none confirms the others.


Production Volume vs. Operational Deployment

The headline figure most cited in humanoid market coverage — 20,000+ units produced in 2025 — conflates three distinct categories: units deployed in real-world operational environments; units used for research, data collection, and training; and units manufactured for entertainment or demonstration. Interact Analysis, in its "Humanoid Robots — 2026" report (May 2026), found that only approximately 10% of 2025 production reached genuine operational environments. That equates to roughly 2,000 units in actual use — a substantial increase from the dozens in deployment in 2024, but far below what production volume headlines imply.


Chinese manufacturers accounted for approximately 80–90% of global 2025 production volume, led by Unitree (~5,500 units sold) and AgiBot (~5,168 units), according to Counterpoint Research data cited by Axis Intelligence (June 2026). The concentration of volume in Chinese platforms built primarily for research and education, not commercial production deployment, explains a significant portion of the gap between production figures and operational deployment counts.


Interact Analysis researcher Blake Griffin stated at the Automate 2026 conference that in 2025 and 2026 there were "virtually no real-world applications" for humanoid robots, with most units deployed for academic research, development, or entertainment in China.


Capital Investment and Market Revenue

Cumulative venture capital investment in humanoid robotics exceeded $9.8 billion by end-2025 (Tracxn data, cited by Axis Intelligence), with Q1 2026 alone adding $2.37 billion — a 288% increase year-on-year. Humanoid robot market revenue in 2025 is estimated at approximately $4.89 billion, with 2026 projected at approximately $6.24 billion (Axis Intelligence cross-source analysis, June 2026).


Market Forecast Context

  • Goldman Sachs (2024, reaffirmed 2025) projects a $38 billion total addressable market for humanoid robots by 2035, revised upward sixfold from a prior $6 billion estimate, driven by a 40% decline in manufacturing costs against their previously forecast 15–20% annual decline rate, and faster-than-expected AI progress ("AI progress surprised us the most"). This is a scenario analysis, not a confirmed market trajectory.

  • Interact Analysis (May 2026) projects market revenue of approximately $15 billion by 2035, with commercial inflection expected in 2032 and annual shipments exceeding 700,000 units by 2035, conditional on achieving economic viability thresholds and embodied AI breakthroughs. The 2.5× divergence between the Goldman Sachs and Interact Analysis 2035 projections reflects fundamentally different assumptions about AI capability timelines — not measurement differences. Both are analyst scenarios.


The International Federation of Robotics opened its first dedicated humanoid robot survey for the World Robotics 2026 report in March 2026. Final IFR humanoid data had not been published as of the research date for this report. IFR 2025 data confirms 542,000 industrial robots installed globally in 2024, with Asia accounting for 74% of new deployments.


Vendor and Platform Landscape

Vendor

Country

Primary Sector

Commercial Status — mid-2026

Funding

Evidence Level

Agility Robotics

USA

Logistics, Manufacturing

9 customer facilities; SPAC announced June 2026 ($2.5B pre-money)

$2.5B SPAC

Level 4

Figure AI

USA

Automotive Manufacturing

Enterprise contracts (BMW, undisclosed logistics customer)

$39B valuation (Series C, Sept 2025)

Level 5

Boston Dynamics

USA

Automotive (Training)

2026 production committed to Hyundai + Google DeepMind

Hyundai-owned

Level 2

Apptronik

USA

Automotive Manufacturing

Active pilot (Mercedes-Benz)

$935M+ Series A

Level 2

Tesla

USA

Internal Only

Internal data collection; no external customers

Public (TSLA)

Level 1/2

Sanctuary AI

Canada

Retail, Automotive

Pilot (Magna International, Mark's/CTC)

~$130M cumulative

Level 2

Unitree

China

Research, Education

Commercially available for purchase (G1, R1)

Shanghai IPO filed

No production deployment

AgiBot

China

Research, Manufacturing

Deployments in Chinese facilities

—

Research/Internal


All funding figures VERIFIED via institutional press releases and reporting. Evidence Levels per the Deployment Evidence Score introduced in Section 3.


Infographic: Humanoid Robot Production vs. Real-World Deployment—2025, showing 20,000+ produced, ~2,000 in use, and 2 verified programs.

3. Deployment Evidence


Which humanoid deployments have actually produced measurable business results?

As of mid-2026, only two enterprise programmes have reported operational outcomes with independent or customer-level confirmation. The Figure AI / BMW Group Spartanburg programme produced independently verified throughput data — 30,000+ vehicles, 90,000+ parts, 1,250+ hours — confirmed by BMW Group communications. The Agility Robotics / GXO Logistics programme produced 100,000+ totes moved in live commercial operation, company-announced and widely corroborated. All other major deployments remain at announcement, pilot, or internal testing stage with no independently published KPIs.


The Deployment Evidence Score

Not all humanoid deployments carry equal evidentiary weight. An announcement that a company has signed a deployment agreement tells a procurement buyer almost nothing about commercial viability. A deployment where a Tier 1 customer has independently confirmed throughput data tells something very specific. The gap between these two data points is, in most cases, the gap between noise and signal in humanoid procurement intelligence.


The Deployment Evidence Score is a five-level classification framework developed by Physical AI Journal to assess the evidentiary quality of every publicly reported humanoid deployment.

Level

Classification

What It Means

Level 1

Announcement Only

A deployment agreement, MOU, or press release has been issued. No operational activity confirmed.

Level 2

Pilot Confirmed

Active testing or pilot deployment at a customer facility is confirmed. No operational KPIs published.

Level 3

Operational Deployment

The robot is performing a defined task in a live production environment. No measurable outcome data published.

Level 4

Measurable Outcomes Published

The deploying company or a corroborating source has published throughput, cycle, or volume data. Customer-independent verification not yet confirmed.

Level 5

Independently Verified Commercial Success

A Tier 1 customer source — company press release, investor filing, or confirmed media statement — has independently confirmed operational KPIs.

Infographic titled Deployment Evidence Score—Physical AI Journal, showing levels 1-5 for humanoid deployments and company examples.

Enterprise Deployment Database

Eight hundred and twenty-six is the number of days between the first confirmed commercial humanoid deployment (GXO Logistics, June 2024) and the research date for this report — the entirety of the commercial humanoid deployment record. The following table classifies every major deployment in that window.

Customer

Vendor

Country

Industry

Task

KPIs Published

Indep. Verification

Level

Confidence

BMW Group

Figure AI

USA

Automotive Mfg

Sheet-metal loading

Yes — 30,000+ vehicles; 90,000+ parts; 1,250+ hrs

BMW Group press release

Level 5

High

GXO Logistics

Agility Robotics

USA

3PL/Logistics

Tote movement (AMR to conveyor)

Yes — 100,000+ totes (co.-announced)

Corroborated, not audited

Level 4

Med-High

Mercedes-Benz

Apptronik

Germany/Hungary

Automotive Mfg

Assembly kit delivery

No

Not available

Level 2

Low

Hyundai RMAC

Boston Dynamics

USA

Automotive (Training)

Parts sequencing (training env.)

No

Company-announced

Level 2

Low

Mark's / CTC

Sanctuary AI

Canada

Retail

110 retail tasks (1-week pilot)

110 tasks — company-claimed

Not independently verified

Level 2

Low

Magna International

Sanctuary AI

Canada

Automotive Mfg

Data collection / task training

No

Partnership confirmed

Level 2

Low

Tesla (Internal)

Tesla

USA

Internal Mfg

Battery cell sorting, data collection

No productivity KPIs

CEO confirmation: learning only

Level 1/2

Very Low

Toyota Motor Mfg Canada

Agility Robotics

Canada

Automotive Mfg

RaaS, 7 units, RAV4 plant

Agreement confirmed Feb 2026

SPAC disclosure

Level 2

Low-Med

Mercado Libre

Agility Robotics

Multiple

E-commerce Logistics

RaaS deployment

Agreement confirmed

SPAC disclosure

Level 2

Low-Med

Schaeffler

Agility Robotics

Multiple

Industrial Mfg

Pilot deployment

Confirmed

SPAC disclosure

Level 2

Low-Med

Evidence Levels assigned per the Deployment Evidence Score. Procurement Confidence reflects evidence quality, not platform capability. This is not an assessment of safety, engineering quality, or technical conformity.


Reading the Evidence Record

The concentration of verified evidence in two deployments — and the absence of independently published KPIs across all others — is the most important pattern in this table for procurement purposes. It does not mean the remaining programmes are failing. It means they cannot currently be used as evidence for procurement business cases.


A commercial agreement between Apptronik and Mercedes-Benz is not equivalent to verified operational performance data. A $935 million fundraise with Mercedes-Benz as a returning investor is compelling capital-market evidence — it is not operational evidence. A buyer who conflates funding confidence with deployment confidence will build a business case on the wrong foundation.


Agility Robotics' SPAC process will require disclosure of audited operational data subject to SEC review — the first time any humanoid robotics company in the US will have its operational claims subjected to public market regulatory scrutiny. The Toyota Motor Manufacturing Canada and Mercado Libre RaaS agreements (both confirmed February 2026) will likely produce Level 3 or Level 4 evidence within 12–18 months of deployment commencement.


4. Economics & ROI Analysis


Can humanoid robots generate a positive ROI today?

Yes — under specific conditions. Structured, repetitive material-handling workflows in controlled environments with high utilisation present a defensible commercial case. Payback periods of 18–36 months are achievable in high-frequency, well-defined deployments per McKinsey Manufacturing Practice benchmarks (2025). ROI depends less on the robot than on task selection, utilisation rate, and accurate total cost of ownership modelling. Outside structured logistics and automotive body-shop work, no verified financial case exists in the public record.


The True Cost Structure

Hardware purchase price represents 55–65% of the true five-year cost of deploying a humanoid robot at enterprise scale, based on procurement data aggregated by the Association for Advancing Automation (A3) in their 2025 Enterprise Robotics Cost Benchmarking Report. The remaining 35–45% divides across four categories.

  • Integration and facility preparation — 10–20% of hardware cost. Covers safety zone commissioning, MES integration, fixture design, and workflow re-engineering. A deployment in an environment already optimised for human-robot collaboration will sit at the lower end of this range; a first-of-kind deployment in a legacy environment will sit at the upper end or beyond it.

  • Maintenance and parts — 10–15% of hardware cost per year. Actuators, sensors, and battery systems are the primary wear items across all current commercial platforms. No commercial humanoid platform has published independently verified actuator replacement schedules from production environments.

  • Software, updates, and licensing — Variable by contract structure. RaaS contracts typically include software updates; capital purchase contracts may not. Confirm software support terms before signing. A robot that cannot receive skill updates in the field will become task-constrained within 12–18 months in a sector where embodied AI models are improving quarterly.

  • Training and supervision — These must include the cost of human supervision required during autonomous operation gaps. After training on over 1 million robot trajectories, autonomous task performance has been independently assessed at approximately 78% — below the 95%+ reliability threshold typically required for unsupervised industrial deployment (Simplexity, December 2025). At that performance level, human fallback coverage for the remaining cycles carries a real cost that must be modelled in every TCO calculation, not treated as a rounding error.


The ROI Decision Framework

ROI viability in humanoid deployment depends on three primary variables. Work through them in sequence.

  • Variable 1 — Task suitability. Humanoid ROI is strongest where the task is repetitive and bounded — same motion, same objects, same spatial geometry — in a controlled environment with mapped layouts, currently staffed by one or more human operators. Tasks with high variability, novel objects, or complex multi-step sequences without a Level 4+ comparable deployment benchmark carry unquantifiable execution risk.

  • Variable 2 — Utilisation rate. A robot running 10-hour shifts five days per week across 48 operational weeks generates approximately 2,400 hours of annual productive time. At double-shift utilisation (4,800 hours/year), payback timelines roughly halve. A robot deployed for a single 8-hour shift generates approximately 1,920 hours per year — and substantially longer payback. The BMW / Figure AI deployment ran 10-hour shifts Monday through Friday (company-claimed, Figure AI). The utilisation design is not incidental to the commercial result.

  • Variable 3 — Integration complexity. A single pick-and-place task at a fixed station in an already-mapped environment is the lowest-integration scenario. Assembly requiring tool changes, unstructured object handling, or dynamic workspace adaptation is the highest. Integration complexity determines both upfront cost and ongoing supervision burden.

Infographic comparing humanoid robot costs, showing hardware is 55–65% of 5-year cost and hidden costs, with bold black text and charts

Buy vs. Lease vs. RaaS

Model

Capital Structure

Maintenance Responsibility

Upgrade Path

Best For

Capital Purchase

High upfront; no recurring hardware fee

Buyer

Buyer-managed

Organisations with internal robotics teams and long-term task confidence

Leasing

Spread over term; residual value consideration

Shared

Defined in lease

Buyers seeking to avoid full capex without full operational flexibility

RaaS (Robots as a Service)

Monthly fee; no hardware purchase

Vendor

Vendor-managed

First deployments; buyers prioritising lower financial risk at entry


Agility Robotics operates a RaaS model. Its investor deck models approximately $8,500/month per robot (company-claimed). The $30/hour figure widely cited in market coverage is the human-labour comparator used in ROI modelling — not the Agility pricing figure. At $8,500/month and two-shift utilisation, break-even against a fully-loaded equivalent human operator position falls in the 12–18 month range before integration, supervision, and downtime costs are applied. The RaaS model shifts maintenance and upgrade risk to the vendor and removes capex from the balance sheet. It also removes hardware ownership and may limit operational flexibility if task requirements change.

For a first humanoid deployment in an unproven task category, RaaS materially reduces financial risk relative to capital purchase.


Pricing Reference

All vendor pricing figures below are company-claimed unless stated otherwise. No independent buyer has publicly confirmed an acquisition cost for enterprise-grade humanoid platforms.

Platform

Published or Target Price

Acquisition Model

Classification

Unitree G1

~$13,500–$16,000 (USD, ex-tax/shipping)

Purchase

Company-claimed (listed price)

Unitree R1 / R1 Air

~$4,290–$5,900

Purchase

Company-claimed

Tesla Optimus

Target $20,000–$30,000 at scale

Not externally available

Company-claimed (Musk, Davos, January 2026)

Agility Digit

~$8,500/month modelled (investor deck)

RaaS

Company-claimed

Figure 03

No list price published

Enterprise contract only

Not publicly available

Boston Dynamics Atlas

Below ~$320,000 (unconfirmed)

Not externally available

Unverified — not confirmed by Boston Dynamics


5. Named Deployment Case Studies


Which humanoid deployments provide the strongest commercial evidence today?

The Figure AI / BMW Group Spartanburg deployment is Level 5 — independently verified by a Tier 1 customer source. The Agility Robotics / GXO Logistics deployment is Level 4 — measurable outcomes published and widely corroborated, though no independent per-cycle cost audit has been released. All other major deployments reviewed here are Level 2: pilots confirmed, operational KPIs not publicly available. The analysis below extracts procurement lessons, not vendor profiles


Case Study 1: Figure AI × BMW Group Plant Spartanburg, South Carolina, USA


Timeline infographic titled Figure AI × BMW Group Spartanburg — Deployment Timeline, showing milestones from 2023 to 2026 and KPI totals.

Deployment Evidence Score:

Level 5 — Independently Verified Commercial Success

Deployment objective: Deploy Figure 02 humanoid robots to perform sheet-metal loading in the body shop at BMW Group's Spartanburg plant — removing sheet-metal parts from racks and placing them on a welding fixture, after which six-axis industrial robots weld and feed parts into the main assembly line.


Implementation timeline: 

Commercial agreement signed November 2023. Robot on-site and testing commenced within six months (company-claimed, Figure AI). Full deployment on the active assembly line achieved within ten months. Total programme duration: 11 months through to Figure 02 retirement in November 2025.


Operational KPIs:

  • 30,000+ BMW X3 vehicles supported — VERIFIED (BMW Group press release)

  • 90,000+ sheet-metal parts loaded — VERIFIED (BMW Group press release)

  • 1,250+ operational hours logged — VERIFIED (BMW Group press release)

  • 10-hour shifts Monday through Friday — VERIFIED (BMW Group press release)

  • 84-second cycle time per part — company-claimed (Figure AI)

  • 37-second load time — company-claimed (Figure AI)

  • >99% successful placement per shift — company-claimed (Figure AI)

  • Target: zero human interventions per shift — company-claimed target; independent intervention rate not published

  • Approximately 1.2 million steps taken (~200 miles) — company-claimed (Figure AI)


Verification assessment: 

BMW Group's own press communications independently confirm the vehicle count, parts count, and operational hours — three figures sourced directly from a Tier 1 customer. This constitutes the strongest independent deployment evidence in the humanoid commercial record as of mid-2026. The KPI figures (placement accuracy, cycle time) originate from Figure AI and are company-claimed. No third-party audit of those metrics has been published.


Deployment limitations: 

Single task, single station, fixed spatial geometry. The BMW deployment does not provide evidence for Figure platforms' performance on diverse manipulation tasks, unstructured environments, or multi-step assembly sequences. Figure 03 has transitioned to a different task at Spartanburg — logistics component sorting — described by Figure AI in June 2026 as a demonstration rather than a fleet-scale commercial rollout.


Procurement lesson: 

Task selection and fixture integration determine success more than robot capability in isolation. The BMW deployment succeeded because the task was chosen to fit the current capability envelope of the platform. Before selecting a platform, identify the task. Verify that a Level 4+ deployment exists in the same task category. Then evaluate the vendor.


Case Study 2: Agility Robotics × GXO Logistics, Flowery Branch, Georgia, USA


Deployment Evidence Score: Level 4 — Measurable Outcomes Published

  • Deployment objective: Deploy Digit humanoid robots to handle tote movement within a live third-party logistics fulfilment environment — moving totes from AMRs to conveyors, stacking containers, and performing last-metre material flow that AMRs cannot execute.

  • Implementation timeline: First commercial deployment at GXO Flowery Branch began June 2024 — the first time a humanoid robot had been deployed at a commercial customer site by any manufacturer (company-claimed, Agility Robotics). The 100,000-tote milestone was announced 20 November 2025. As of the Agility Robotics SPAC disclosure (June 2026), Digit has accumulated over 65,000 hours of operation across nine customer facilities (company-claimed).


Operational KPIs:

  • 100,000+ totes moved at GXO Flowery Branch — company-announced (November 2025); corroborated by Robotics & Automation News, Assembly Magazine, and Interesting Engineering

  • 65,000+ total commercial operating hours across nine facilities — company-claimed (Agility SPAC disclosure, June 2026; subject to SEC review)

  • $300 million+ in booked multi-year Digit v5 orders — company-claimed (Agility SPAC disclosure; subject to contractual milestones)


Peggy Johnson, CEO of Agility Robotics, stated at the SPAC announcement on 24 June 2026:

"Humanoid robots are a critical driver of American technology leadership and the future of global industry. With category-defining commercially deployed humanoid robots operating in real customer environments today, Agility is at the forefront of a new era where safety-first, AI-powered technology can reliably work alongside people."

(TechCrunch, June 2026)


Verification assessment: 

The 100,000-tote figure is company-announced and has not been independently audited. It has been corroborated across multiple trade press sources without contradiction. GXO Logistics has not published an independent statement of operational KPIs. Intervention rate, downtime hours, and cost per tote versus human operator have not been disclosed publicly.


Deployment limitations: 

Tote-handling in a structured warehouse is the task category most favourable to current humanoid capability. This evidence does not extend to picking of varied SKUs, unstructured object handling, or high-dexterity assembly tasks. The absence of per-cycle cost data prevents independent ROI confirmation.


Procurement lesson: 

Volume throughput across months of commercial operation is meaningful evidence. It is not the same as a cost-per-cycle analysis. Before approving a comparable warehouse deployment, request: (1) operational hours per robot per shift; (2) human interventions per 1,000 cycles; (3) unplanned downtime hours; and (4) total cost including integration, supervision, and maintenance.


Case Study 3: Apptronik × Mercedes-Benz, Berlin and Kecskemét


Deployment Evidence Score: Level 2 — Pilot Confirmed

Deployment objective: 

Deploy Apollo humanoid robots at Mercedes-Benz manufacturing facilities for intra-logistics tasks — delivering assembly kits, moving parts to production lines, and conducting initial component inspections.


Implementation timeline: 

Commercial agreement announced 15 March 2024. Active pilots at the Berlin Digital Factory Campus (Marienfelde) and Kecskemét plant (Hungary) confirmed through 2025–2026. Apptronik closed a $520 million Series A extension in February 2026 — bringing total Series A funding to over $935 million at a valuation of approximately $5.5 billion (Bloomberg) — with Mercedes-Benz participating as a returning investor.


Operational KPIs:

 None published as of mid-2026. Mercedes-Benz has not released throughput, task completion, or productivity comparison data. Apptronik has not published operational metrics from the deployment.


Verification assessment: 

The commercial agreement is VERIFIED (PR Newswire, March 2024). The ongoing partnership is VERIFIED via Mercedes-Benz investor participation in the February 2026 round (Bloomberg). No operational outcome data has been independently confirmed.


Procurement lesson: Investor confidence in a deployment partnership is not a substitute for operational outcome data. A $935 million fundraise with a major OEM as a returning investor demonstrates long-term commercial confidence. It tells a procurement buyer nothing about what the robot has delivered in operation. Evaluate Level 2 programmes as "watch" positions — not as capital-commitment evidence.


Case Study 4: Boston Dynamics Atlas × Hyundai Robotics Metaplant Application Centre, USA


Deployment Evidence Score: Level 2 — Training Environment Confirmed

Deployment objective: 

Deploy the production version of Atlas at Hyundai's RMAC — described by Boston Dynamics CEO Robert Playter as a "data factory" (AI2Work, May 2026) — to build a training dataset for humanoid manufacturing skills, beginning with parts sequencing in automotive contexts.


Implementation timeline: 

Production Atlas unveiled at CES 2026 (January 5, 2026). Boston Dynamics announced immediate production at its Boston headquarters. Entire 2026 Atlas production run committed to Hyundai RMAC and Google DeepMind — no external customers until 2027. Full Hyundai factory deployment at Hyundai Metaplant America (Savannah, Georgia) targeted 2028 (company-claimed).


Operational KPIs: 

None published. RMAC is explicitly a controlled training and research environment, not a production line generating throughput evidence.


Verification assessment: 

The CES 2026 announcement is VERIFIED (Boston Dynamics press release, Hyundai Motor Group newsroom). The RMAC deployment as a training environment is VERIFIED. The 2028 full-factory timeline is company-claimed.


Labour context note: 

The Korean Metal Workers' Union issued a statement in January 2026 opposing Atlas deployment at Hyundai factories without a labour-management agreement, citing the robot's anticipated cost relative to worker wages. This is a factual matter of record for procurement teams evaluating deployment in environments with organised labour.


Procurement lesson: 

Distinguish training environment deployments from production deployments. The RMAC is designed to generate training data, not throughput. Evidence produced there will eventually inform production readiness evaluations — but it is not production evidence today. Procurement teams evaluating Atlas for 2026–2027 deployment should plan against Level 2 evidence, with production-level evidence not expected before 2027 at the earliest.


6. Friction, Risk & Unresolved Issues


What are the biggest risks of deploying humanoid robots today?

The primary enterprise risks stem from operational uncertainty, not hardware capability alone. The six risk categories below are framed exclusively as procurement risks — factors that materially affect deployment feasibility, total cost, and business case confidence. None are resolved by demonstration footage, funding rounds, or upward forecast revisions. Each requires a direct contractual or operational response before capital is committed.


Risk 1: Safety Certification Vacuum

ISO 25785-1 — the only international safety standard specifically addressing dynamically stable humanoid robots — remains at Working Draft stage as of mid-2026. The standard was approved by ISO Technical Committee 299 for development; the US delegation includes representatives from Agility


Robotics, Boston Dynamics, and the Association for Advancing Automation (A3). No confirmed publication date exists. Industry estimates put the timeline at 18–36 months from the May 2025 working draft to a ratified standard. Current deployments operate under ISO 10218-1:2025 and ISO 10218-2:2025, which were designed for statically stable robots and do not directly address fall risk — a primary hazard specific to walking bipedal systems.


Procurement impact: 

Buyers cannot obtain third-party certification of a humanoid deployment under a published humanoid-specific safety standard. Risk assessment, facility preparation, and SLA terms must be negotiated on the basis of manufacturer guidance and ISO 10218 principles applied by analogy. Build this gap explicitly into any risk register.


Risk 2: The Teleoperation–Autonomy Gap

Humanoid robots in commercial deployments in 2024–2026 are not fully autonomous in the manner implied by much vendor communication. Teleoperation is a standard component of current deployment operations — a legitimate data-collection methodology, but not a scalable commercial model. After training on over 1 million robot trajectories across 217 distinct tasks, autonomous task performance in one independent technical assessment improved to approximately 78% — still below the 95%+ reliability threshold typically required for unsupervised industrial deployment (Simplexity, December 2025).


Procurement impact: 

Human supervision and fallback intervention capability must be included in total cost of ownership. Request the vendor's contractually specified autonomous operation rate and the precise definition of "intervention" in SLA terms. A deployment achieving 78% autonomous performance still requires human coverage for the remaining 22% of cycles — and that coverage has a real cost.


Risk 3: No Published Uptime Benchmarks

Conventional fixed industrial robots achieve 95–99% uptime in well-maintained environments. No commercial humanoid platform has published independently verified uptime data at a comparable standard as of mid-2026. The mechanical complexity of a bipedal humanoid — hundreds of joints, actuators, and sensors, each a potential failure point — creates a substantially different reliability profile from a six-axis industrial arm. Battery life ranges from 2–8 hours per charge across current platforms; charging intervals disrupt continuous shift coverage.


Procurement impact: 

Shift planning assumptions cannot be grounded in published, independently verified uptime data. Require contractually binding uptime guarantees in SLA terms, and treat any guarantee unsupported by operational data from a comparable deployment with substantial caution.


Risk 4: Total Cost Substantially Exceeds Hardware Price

Hardware represents 55–65% of the true five-year cost of enterprise humanoid deployment (A3 Enterprise Robotics Cost Benchmarking Report, 2025). The remaining 35–45% covers integration, maintenance, software, training, supervision, safety infrastructure, and insurance.


Procurement impact: 

Business cases built on robot unit price produce systematically underestimated cost forecasts. A robot acquired at any price point carries total five-year costs of approximately 1.5×–1.8× the hardware acquisition cost once integration, maintenance, supervision, and safety infrastructure are included — per the A3 benchmark. Require a full five-year TCO model before approving any pilot budget.


Risk 5: Narrow Task Envelope

All commercially validated humanoid deployments as of mid-2026 are confined to structured, repetitive material handling: tote movement at fixed AMR handoff points; sheet-metal loading at defined welding fixtures; parts sequencing in controlled cells. No commercially verified deployment has demonstrated reliable autonomous performance across a diverse task set, in an unstructured environment, or in any workflow requiring multi-step manipulation of novel objects.


Procurement impact: 

Vendor claims about general-purpose capability cannot be evaluated against current commercial evidence. Scope any initial deployment to a task within the validated evidence envelope. Treat general-purpose capability claims as forward-looking aspirations, not current operational evidence.


Risk 6: Production Timeline History — Tesla Optimus

Every Tesla Optimus production target since 2022 has been missed. The 2025 target of 5,000–10,000 productive units was confirmed unmet on the Q4 2025 earnings call (January 2026), when


Elon Musk, CEO of Tesla, stated (Q4 2025 Earnings Call, 28 January 2026):

"We have several hundred units deployed, primarily for learning, not productive tasks — still very much in the R&D phase."

The 2025 actual output was confirmed as "several hundred" units — a miss of over 80% against the stated target.


Procurement impact: 

Buyers whose deployment plans depend on Tesla Optimus external availability timelines — currently stated as late 2027 (company-claimed) — should apply substantial uncertainty margins. Build alternative platform options into any procurement plan that includes Optimus on its critical path.


7. Competitive Platform Comparison


Which humanoid platform currently has the strongest enterprise deployment evidence?

Agility Robotics Digit holds the deepest commercial deployment record by operating hours — 65,000+ across 9 facilities (company-claimed, SPAC disclosure, June 2026). Figure AI Figure 02/03 holds the only Level 5 independently verified production-line result, confirmed by BMW Group. Boston Dynamics Atlas and Apptronik Apollo hold Level 2 evidence. Tesla Optimus is Level 1/2 — internal only. Unitree G1/R1 is commercially available from approximately $13,500 (company-listed) with no documented commercial production deployment.


The Platform Readiness Index

The Platform Readiness Index evaluates platforms on procurement-relevant evidence criteria, not technical specifications. Specifications describe what a robot is built to do in controlled conditions. Evidence describes what it has demonstrably delivered in commercial operation.

Vendor

Platform

Evidence Level

Verified KPIs

Commercial Availability

Pricing Transparency

Procurement Readiness

Figure AI

Figure 03

Level 5

Yes — BMW Group (Tier 1)

Enterprise contract only

Not published

Proven (automotive body-shop)

Agility Robotics

Digit

Level 4

Corroborated, not audited

RaaS; no purchase

~$8,500/mo modelled (co.-claimed)

Early Production Evidence (logistics)

Apptronik

Apollo

Level 2

No

Enterprise pilot

Not published

Pilot Stage

Boston Dynamics

Atlas

Level 2

No

Committed to Hyundai + DeepMind; ext. customers 2027

Not published

Training/Data Stage

Tesla

Optimus Gen 3

Level 1/2

No

Not externally available

Target $20K–$30K (co.-claimed)

Internal Only

Sanctuary AI

Phoenix

Level 2

110 tasks/1-wk pilot (co.-claimed)

Limited pilot

Not published

Pilot Stage

Unitree

G1, R1

Not rated

No commercial production deployment

Available from ~$13,500 (listed)

Transparent

Research/Education Only

Procurement Readiness categories: Proven Commercial Deployment | Early Production Evidence | Pilot Stage | Training/Data Stage | Internal Only. This index is a procurement evidence assessment — not an evaluation of engineering quality, safety, or technical conformity.


Platform Readiness Index bar chart: Figure AI 5.0 leads; Agility 4.0, others lower. Bottom callouts show 10% deployment and ISO draft.

Interpreting the Index Correctly

Evidence absence and task mismatch are the two most common errors in applying platform comparison data to procurement decisions.


Task fit overrides evidence score. Agility Robotics Digit's Level 4 evidence is in logistics tote-handling. Figure AI's Level 5 evidence is in automotive body-shop sheet-metal loading. A buyer seeking a logistics deployment should weight Digit's evidence more heavily for that task, even though both platforms carry high evidence scores in their respective domains. Evidence in one task category does not transfer automatically to a different task category.


Evidence absence is not platform failure. Apptronik Apollo's Level 2 rating reflects the absence of published operational data — not a finding about the platform's capability or the Mercedes programme's performance. A Level 2 rating means: "Sufficient evidence to make a procurement recommendation in this category does not yet exist."


8. Strategic Recommendations


Should enterprises invest in humanoid robots now or wait?

The answer is task-specific and evidence-gated — not technology-gated. For structured, repetitive material-handling tasks in controlled environments where at least a Level 4 deployment exists in the same task category, a pilot investment is defensible. For all other task types and environments, the appropriate position is structured market monitoring. The decision framework below identifies which of those positions applies to a specific deployment proposal.


Deploy Now

Logistics tote-handling and automotive body-shop sheet-metal loading are the only two task categories with Level 4+ deployment evidence as of mid-2026 — the boundary that determines where a pilot investment is currently defensible. A pilot is justifiable when all of the following conditions are simultaneously satisfied:


1. The task is structured and repetitive — defined cycle time, known objects, fixed spatial geometry.

2. The environment is controlled — pre-mapped, stable, with accessible safety zone commissioning.

3. High utilisation is achievable — the task runs for at least two shifts per day, five or more days per week. Single-shift utilisation materially extends payback timelines.

4. An integration partner with prior humanoid experience is available and contractually engaged.

5. A Level 4 or Level 5 deployment in the same task category exists in the public record.

6. Total cost of ownership has been modelled using the A3 benchmark — hardware at 55–65% of five-year cost, with integration, maintenance, supervision, and safety infrastructure fully accounted for.


Observe

The 12–18 months following June 2026 will likely produce Level 3 or Level 4 evidence from several programmes currently at Level 2 — including Toyota Motor Manufacturing Canada (Agility Robotics, February 2026 agreement, 7 units confirmed) and Mercado Libre (Agility Robotics, February 2026 agreement). Structured market monitoring is appropriate when:

  • Tasks require flexible manipulation or varied object handling.

  • Environments need extensive facility modification before commissioning.

  • The closest comparable deployment is Level 2 or below.

  • The organisation lacks internal robotics integration capability or a qualified integration partner.


Wait

Capital commitment is premature for the following categories — no commercially verified deployment evidence exists as of mid-2026:

  • Healthcare — no commercially deployed humanoid in a clinical or care environment with published operational KPIs.

  • Construction — no commercially deployed humanoid on an active construction site with published KPIs.

  • General-purpose retail — the Sanctuary AI / Mark's pilot (Level 2, one week, company-claimed) does not constitute evidence for commercial retail deployment.

  • Hospitality — no Level 3+ deployment in the public record.

  • Home use — no platform commercially available for unsupervised home operation at industrial reliability standards.

Infographic titled Humanoid Robot Deployment Decision Matrix — 2026, with Deploy Now, Observe, and Wait panels and task examples.

The Procurement Evidence Checklist

Before approving any humanoid deployment, require satisfactory answers to all seven of the following questions. A vendor that cannot answer all seven has not completed the preparatory work required to justify an enterprise capital commitment.

  1. Reference evidence: Has a reference customer in the same task category published independently verifiable KPIs? What is the Evidence Score of the closest comparable deployment?

  2. Full cost model: Is integration complexity fully costed — not just hardware? Has a five-year TCO been constructed using the A3 benchmark?

  3. Supervision cost: What human supervision or fallback labour is required at the vendor's stated autonomous operation rate? Is this cost included in the business case?

  4. SLA terms: What is the vendor's stated uptime guarantee? Is it contractually binding with defined remedies for underperformance?

  5. Exit strategy: What is the contractual exit plan if the deployment underperforms its KPI commitments at the scheduled review point?

  6. Safety framework: What safety assessment has been completed, under which standards, and by what body? Note that ISO 25785-1 (humanoid-specific) has not been published as of mid-2026.

  7. Timeline risk: Does the deployment timeline account for potential delays in software maturity, integration complexity, and regulatory development?


9. Executive FAQ


FAQ 1: Have humanoid robots demonstrated measurable ROI in real enterprise deployments?

Yes — in two documented cases. The BMW Group / Figure AI deployment at Plant Spartanburg (Level 5) produced 30,000+ vehicles supported, 90,000+ parts handled, and 1,250+ operational hours, confirmed directly by BMW Group press communications; the GXO Logistics / Agility Robotics deployment (Level 4) moved more than 100,000 totes in live commercial operation, company-announced and widely corroborated. Neither deployment has published independent cost-per-cycle data, so ROI as a financial metric has not been verified for either programme. A vendor claiming ROI based on a deployment in a different task category is not providing relevant evidence for your use case.


FAQ 2: Which humanoid robot has the strongest enterprise deployment evidence today?

Agility Robotics Digit holds the deepest commercial deployment record by operating hours — 65,000+ across 9 facilities (company-claimed, SPAC disclosure, June 2026); Figure AI Figure 02/03 holds the only Level 5 independently verified production-line result, confirmed by BMW Group. Both findings are task-specific: Digit's evidence is in logistics tote-handling; Figure's is in automotive body-shop sheet-metal loading. Boston Dynamics Atlas, Apptronik Apollo, and Tesla Optimus all hold Level 2 evidence — confirmed pilots or agreements, no independently published KPIs. Task fit is as important as evidence score; the platform with the strongest evidence for your specific task is the starting point.


FAQ 3: Which industrial tasks are humanoid robots consistently succeeding at?

Two task categories have commercially validated evidence as of mid-2026: logistics tote-handling (Agility Robotics Digit at GXO Logistics, Level 4, 100,000+ totes) and automotive body-shop sheet-metal loading (Figure AI Figure 02 at BMW Group Plant Spartanburg, Level 5, 30,000+ vehicles, 1,250+ hours). All other tasks — complex assembly, diverse SKU picking, machine tending with varied geometries, healthcare, construction — have no comparable deployment evidence in the public record. Define the specific task before selecting a platform; if it falls outside these two categories, treat the deployment as a first-of-kind pilot with contractual protections calibrated accordingly.


FAQ 4: Should manufacturers launch a humanoid pilot in 2026 — or wait?

Deploy now if the task is structured and repetitive, a Level 4+ deployment exists in the same task category, utilisation can reach double-shift levels, and total cost has been modelled against the A3 benchmark. Observe and monitor if the task requires flexible manipulation, the closest comparable deployment is Level 2, or no integration partner with humanoid experience is available. Wait if the target is healthcare, construction, general-purpose retail, or any task without a Level 3+ comparable deployment in the public record. Apply the Procurement Evidence Checklist in Section 8 to any specific proposal before it reaches capital committee.


FAQ 5: What evidence should procurement teams require before approving a humanoid deployment?

Require satisfactory answers to all seven criteria in the Procurement Evidence Checklist™: a reference customer in the same task category with independently published KPIs, Evidence Score of the closest comparable deployment, full five-year TCO model, defined supervision costs, contractually binding uptime SLA with remedies, defined exit strategy, and a written safety assessment noting the ISO 25785-1 gap. No vendor demonstration or funding announcement satisfies any of these seven requirements. Require contractual provisions for key performance milestones at six months and twelve months from deployment commencement.


FAQ 6: How should enterprises compare humanoid vendors beyond technical specifications?

Compare vendors on six criteria in priority order: Evidence Score in the same task category; commercial availability; deployment sector match; integration ecosystem maturity; pricing transparency; and repeat-customer validation. Technical specifications — payload, speed, degrees of freedom, battery life — are company-claimed figures describing controlled-condition capability, not verified commercial operating performance. A vendor with Level 4 evidence and limited spec disclosure is a more defensible procurement choice than a vendor with published specs and Level 1 evidence. Score your candidate vendors against these six criteria and the Procurement Evidence Checklist in Section 8 before any capital commitment is approved — that evaluation, not vendor demonstrations, is the basis for a defensible procurement decision in 2026.


10. Scope & Disclaimer

This report is published by Physical AI Journal, an independent research publication operated by Sekason Research Limited, London, United Kingdom (Company Registration Number: 14339910).

What this report covers: Market and adoption intelligence on humanoid robot deployments (2024–2026). Procurement decision frameworks based on publicly available evidence. Classification of deployment evidence by quality and independence. Analysis of economic and commercial factors relevant to automation investment decisions.


What this report does not cover: This report does not constitute legal, financial, investment, engineering, safety-certification, or professional advice of any kind. Physical AI Journal does not evaluate, certify, or endorse the safety, technical conformity, or fitness-for-purpose of any robotic platform, product, or deployment. No assessment of engineering quality, safety design, or technical conformity of any specific platform is stated or implied anywhere in this report.


Company-claimed figures are identified as such throughout this report and have not been independently verified by Physical AI Journal or Sekason Research Limited. Readers should verify any such figures directly with the relevant manufacturer or vendor before making decisions based on them. Platform specifications, pricing, and deployment status in the physical AI and humanoid robotics sector may change rapidly and without notice.

Readers should obtain independent professional, technical, and legal advice before making procurement, deployment, safety-certification, or investment decisions.

Full terms and disclaimer: physicalaijournal.org/disclaimer


Physical AI Journal Research Team — Independent analysis by Sekason Research Limited, United Kingdom

References and Sources

This report is supported by a combination of authoritative industry research, company disclosures, independent journalism, analyst reports, and recognised robotics organisations. Every source has been classified according to its reliability (Tier 1–4) and whether the information is VERIFIED through independent or official sources, or COMPANY-CLAIMED where the data originates from vendor announcements or press releases.



This report is backed by authoritative research, independent verification, and a structured analytical methodology. All numerical data and deployment outcomes have been clearly classified as either independently VERIFIED or COMPANY-CLAIMED to ensure transparency and support informed enterprise decision-making.


Physical AI Journal Research Team — Independent analysis by Sekason Research Limited, United Kingdom

bottom of page