July 22, 2026

A simple recipe for winning robotics

Most robotics companies are trying to perfect manipulation and locomotion before they deploy. That forces them to overcome a brutal reliability cliff before they can scale.We believe that’s a strategic error. Start with perception. Collapse deployment cost. Capture territory. Enable co-embodiment.

Robots can be said to have capabilities that roughly map into three categories:

Locomotiontraversing environments, typically to carry a payload.

Manipulationaffecting objects in the environment.

Perceptionmaking observations about the environment.

Modern robotics combine these capabilities, allowing them to perform more complex tasks in challenging and dynamic environments, but for the sake of clarity, let us imagine the simplest and purest instantiations of each:

🦶 A simple locomotion robot will move from A to B on command, without needing to perceive the world or affect its surroundings. See: an elevator.

✋A simple manipulation robot will affect an object in the environment without having to perceive the object, or having to move to find the object. See: a mechanized hammer.

👁️ A simple perception robot will need to observe the environment, without having to manipulate the environment or transport itself to perform its task: See:  a people-counting CCTV camera.

Correspondingly, we can think of categories of tasks in the world that map quite well to these capabilities: there are locomotion tasks, manipulation tasks, and perception tasks.

Tesla cars are locomotion robots.

Weave laundry folders are manipulation robots.

Kabam mall security rovers are perception robots.

Tesla cars obviously rely heavily on perception, but they do so in order to perform a locomotion task, just as Weave’s robots rely on perception in order to perform a manipulation task, and Kabam relies on locomotion capabilities in order to perform a perception task.

Today, almost all capital and attention in robotics is going towards perfecting robots for locomotion and manipulation tasks - but I want to show you that this is almost certainly a strategic error.

Deployment Considerations

When considering a robot for a task, there are three helpful dimensions of analysis to help us determine if the robot is right for the job:

Value - Can the robot perform the job cheaper or better?

Cost - How much will it cost to develop, deploy and maintain the robot?

Risk - How disruptive would it be if the robot malfunctioned or failed? How much uptime can you guarantee, and how catastrophic are the failure modes?

The challenge in finding good targets for deployment is finding where you can either perform a job cheaper or better, where the cost of getting started isn’t prohibitive, and the risk of failure isn’t too costly.

For a manipulation example, it is very costly when a robot on an auto assembly line malfunctions, because the car might be damaged - or the whole assembly line might stop until the issue has been addressed. The riskier the environment is, the more 9s of reliability you have to be able to guarantee the customer to get your foot in the door and to scale.

Looking at value and cost together gives us a sense of time to ROI, and risk gives us a sense of how much reliability we need before the deployment is viable at all. There is often a brutal reliability cliff that must be overcome before the robots can come out of pilot stage.

Looking at locomotion, self-driving cars produce valuable work, and they need not cost much more than a human-operated vehicle - but the failure modes are incredibly harmful.

Surprisingly, it turns out that there are many perception tasks in the world that are more valuable than commonly pursued manipulation tasks, with a much lower cost of R&D and deployment - but also with comparatively trivial risk profiles.

Example: Robots for Vineyards

My good friend @jmoonio founded Budbreak, an agricultural robotics company, based on this perception-first logic. Naively, if you say you’re working on robots for vineyards many will assume that you’re looking to solve a manipulation task like harvesting, pruning or spraying. However, it would be extremely difficult today to do either of those tasks better than a human or cheaper than minimum wage. It would take an extraordinary effort to develop something that could solve those tasks 10% better or cheaper.

But a single crop-destroying disease destroys up to 60,000 dollars worth of grapes per hectare every year, and by building rugged “Mars rovers for vineyards”, Budbreak helps the vineyard increase their yield rather than reduce their costs. As a deployment strategy, this is a combination of high value, low cost and low risk. There are tens of millions of addressable permanent crop acres in the US alone, and billions to be made securing the food supply through perception alone.

Giving the rover a modular arm, initially used to bring a camera closer for inspection, means that they have a robot mechanically capable of addressing manipulation tasks in the future - when they’ve already captured that territory.

Example: Robots for Retail

More than robots that can help with restocking, retailers desperately need robots to make the most of the staff, shelves and resources that they have - and for that they need perception, rather than manipulation.

Late last year we were visited by a large European retailer, who told us that they estimate that they lose 30-60 minutes per employee per day across their 10,000+ associates, predominantly in the morning, due to the staff not really being sure what to do. Rather than a robot to replace their staffers, their ask was simple:

Can you make us a robot that drives around the store at night to figure out what needs get to done?

We had already learned with a previous retailer that a “simple” task manager in shared augmented reality allowed for better asynchronous knowledge and task transfers, saving 15 minutes per employee per day, and now this retailer wanted to take a step further and have AI populate the task list.

Another customer of ours expanded the problem-statement, explaining that tough competition means their stores have to be open from early in the morning to late at night, and that they only really have a store manager present for some of those hours. With high staff turnover, and a workforce that is often low skill and poorly motivated, those unsupervised hours are a constant drain on the business.

Another retailer added the most surprising dimension of the problem-statement: emotions. They struggle with retaining staff, which keeps training costs perpetually high, in part due to the natural friction that comes from boring tasks being assigned by another human. Their hope is that AI-assigned tasks will make the workplace less toxic, more enjoyable, and that staff will retain longer.

More than robots that can help with restocking, these retailers desperately need robots to make the most of the staff, shelves and resources that they have - and for that they need perception, rather than manipulation.

This year we plan to deploy wheeled humanoids to these domains to help with task generation, empty shelf detection, compliance checks and data capture, augmenting the value that our customers are already getting from the AI copilots we provide for their handhelds.

As with the vineyard example, this is a deployment target with high value, lower cost and much lower risk.

The Value of Deployment

There are many opportunities today to start massively scalable deployments of perception robots, and early movers here will gain compounding benefits that may allow them to dominate better funded labs that are going for manipulation and locomotion.

Deploying gives you many benefits as a physical AI company:

  • Experience - Operational know-how and muscle memory
  • Revenue - Funds to accelerate your mission
  • Supply chain - Iteratively design and discover better ways of shipping
  • Data - Capture real-world data and create potential flywheels
  • Territory - Capture and control scarce deployment opportunities

We can imagine a world where robots themselves become critical infrastructure, offering co-embodiment to multiple intelligence providers.

If you deploy hardware that already solves valuable perception tasks, but is also mechanically capable of manipulation given better software, you create gravitational pull. Intelligence providers are incentivized, or even compelled, to deploy onto the robots already in the field.

This is robotics as infrastructure: a mechanical substrate for AI to inhabit. An internet of sensing and actuation.

The path to zero cost

By building protocols for collaborative perception, and treating the phones and glasses as almost-robots in their own right, we made it possible to pre-deploy the environments before the actual robots arrive.

@JacklouisP recently noted that while the time required for each robot deployment has gone from 4-6 weeks down to 1-2 weeks for best in class projects, the ultimate goal is to reduce that to zero. “If your roadmap doesn't include this KPI, you have a demo reel, not a deployment strategy”.

Earlier this year, @oyhsu of a16z made a similar point, calling for better robotics middleware and deployment automation to bring down the cost of deployment to something that can realistically scale.

When I first met @jacobrintamaki, one of his first questions was “what is your non-consensus bet on how to source forward deployed engineers? If you are to deploy to 100,000 locations, will that not require thousands of FDEs?”

Again, perception comes to the rescue. More specifically, collaborative perception.

We started building our platform for collaborative machine perception in 2021, based on belief that you could get a head start on robotics by treating smartphones and smart glasses as robots with no arms and legs.

The minimum viable form factor for performing meaningful perception tasks was already in our users’ pockets, and we would leverage that to start building a moat long before the hardware for locomotion was ready.

Start with perception. In fact, start with phones to pre-deploy the environment and give immediate utility. Radically reduce the cost of deployment. Capture territory. Offer up the extensible robot already in the domain as a platform for other providers. Embrace co-embodiment.

We built an AI copilot for retailers that allowed them to create digital twins of their environments, making them browsable, searchable and navigable. This allowed us to give the retailers new data about their stores, AI-generated tasks and recommendations, AR navigation for staff and customers, and a powerful AR task manager that allowed for asynchronous communication (and, eventually, communication with the robots.)

Importantly, we made this copilot simple enough to deploy that the retailers can do it themselves, without us sending FDEs to the thousands of locations that have already signed up to make their environments AI accessible.

A few months ago, we demonstrated how environments already set up for use with the phones, by the staffers themselves, was made instantly accessibly to robots. By just showing the robot a QR code it would identify the customer site, how to contact their cloud and edge services, receive the relevant spatial context, and be ready for simple value generating tasks.

By building protocols for collaborative perception, and treating the phones and glasses as almost-robots in their own right, we made it possible to pre-deploy the environments before the actual robots arrive.

The Case for Co-Embodiment

Despite software having eaten the world, roughly 70% of the world economy is still tied to physical locations and physical labor, measured in both headcount and GDP. Therefore, the biggest opportunity for AI to be transformative is in the physical world.

For this to happen, I submit to you that the internet has to grow in three new dimensions:

  • An internet of spaces
  • An internet of sensors
  • An internet of actuators

The real world web is our attempt to make the physical world accessible to AI, making deployments easier and collaboration smoother by providing an early internet of spaces and sensors. Multiple devices and robots can collaboratively map and reason about a domain in a privacy-preserving way, with granular role-based access controls and the ability to self-host while still being discoverable and interoperable with third parties.

What comes next, we imagine, is more advanced co-embodiment inside the robots themselves: the internet of actuators.

In a world where robots like the ones deployed by Auki, Budbreak and other real world web developers have already captured territory, there is ample incentive for other intelligence providers to focus their efforts on that platform to reach customers at much lower cost to themselves and the customer.

Former COO of K-Scale, Rui Xu, has written about AI-native hardware, and how we can re-imagine how the devices around us present themselves to our new agents and companions. Ultimately, I think the logical conclusion of Rui’s vision is to include the robots themselves as part of this infrastructure layer.

When hardware becomes a Tool, the AI becomes the craftsman, and the possibilities become infinite. - Rui Xu

This is by no means how robots work today. Both cross-embodiment and co-embodiment are largely unsolved problems, but if we want to accelerate the benefits of physical AI I think it is important that we push the envelope of what is possible in this regard and embrace co-embodiment as a shared go-to-market for the best intelligence providers and hardware OEMs in the world.

The open source operating system for smart glasses, MentraOS, has already started embracing this as they design the next inevitable hardware platform. The founders of Mentra made the wise decision to focus on an open source platform based on a unique realization:

The multitude of AI agents that we wish to bring with us out into the real world will need shared access to the context provided by the glasses. They need shared access to camera, the microphone, the IMU, but also to the display and speakers, so that they can be proactive and hands-free.

Watch: MentraOS at AWE2025

Rather than trying to build one AI to rule them all, Mentra embraced the idea that different app developers and intelligence providers will excel in different areas, and that the end user will want to pick and choose which utilities they run together at the same time. I think the robotics industry would be wise to embrace this approach, as well. In a world where multiple applications and agents can share access to the robot-as-infrastructure, both the end users and developer ecosystem will benefit.

The Bull Case for Capturing Territory

There are millions of retailers in the world today that manage complex physical environments with very little assistance from AI. After all, how can AI assist them when the physical world is beyond their reach?

Tens of millions of people work on retail floors, helping customers, moving products, stocking shelves and tidying up. They’re often lower on the skill and education spectrum, staff turnover is incredibly high, and language difficulties are common in many modern multi-cultural countries.

Through a combination of phones, glasses and robots, these environments can be made more productive, profitable, but also nicer places to work and visit.

Last summer we released our first co-pilot for retail workers, and the results have been impressive:

  • 15 minutes saved per employee per day on handovers
  • 20-40% reduced walking distance for staff
  • Faster order fulfillment
  • Lower training cost

The retailers are already paying us 500 dollars per location per month for the phone-based copilot - comparable to what early manipulation robots like Weave and 1x capture, but without the complexity and scaling bottlenecks.

This year, we’re rolling out task-generating glasses, performing perception tasks while the staffers are on the move, and robots that audit the state of the store and continuously and autonomously capture more data that helps the stores run better.

We believe we can realistically scale this into the order of 100,000 retail locations by the end of 2030, because our customers themselves are the FDEs, and the utility is undeniable.

100,000 robots generating 100 dollars of day per revenue, and a million glasses generating 10 dollars per day revenue, and 500 dollars per location per month just for the phone companion means you can realistically scale just retail perception to many billions in recurring revenue in a relatively short timeframe.

These 100,000 robots and 1,000,000 million glasses and countless smartphones will be the mechanical substrate for increasingly intelligent agents to inhabit - the next extension of the internet itself.

But retail and agriculture is just the beginning. We’re building the decentralized nervous system for physical AI, so that 70% of the world economy can be brought into the light of the AI era.

Authored by Nils Pihl, CEO @ Auki Labs

アウキ・ラボについて

Aukiはポーズメッシュという地球上、そしてその先の1000億の人々、デバイス、AIのための分散型機械認識ネットワークを構築しています。ポーズメッシュは、機械やAIが物理的世界を理解するために使用可能な、外部的かつ協調的な空間感覚です。

私たちの使命は、人々の相互認知能力、つまり私たちが互いに、そしてAIとともに考え、経験し、問題を解決する能力を向上させることです。人間の能力を拡大させる最も良い方法は、他者と協力することです。私たちは、意識を拡張するテクノロジーを構築し、コミュニケーションの摩擦を減らし、心の橋渡しをします。

ポーズメッシュについて

ポーズメッシュは、分散型で、ブロックチェーンベースの空間コンピューティングネットワークを動かすオープンソースのプロトコルです。

ポーズメッシュは、空間コンピューティングが協調的でプライバシーを保護する未来をもたらすよう設計されています。いかなる組織の監視能力も制限し、空間のプライベートな地図の自己所有権を奨励します。

分散化はまた、特に低レイテンシが重要な共同ARセッションにおいて、競争優位性を有します。ポスメッシュは分散化運動の次のステップであり、成長するテック大手のパワーに対抗するものです。

アウキ・ラボはポスメッシュにより、ポーズメッシュのソフトウェア・インフラの開発を託されました。

Keep up with Auki

ニュース、写真、イベント、ビジネスの最新情報をお届けします。

最新情報