Back to Perspectives
August 16, 2026
Physical AI & Robotics

One Line in Sahil Bhaiwala's Data Thesis Points at the Next $100B Question

Handshake AI's Sahil Bhaiwala made the sharpest case I've read this year for the data-authoring economy, and scoped it carefully enough to set physical intelligence aside for a future post. That's exactly the layer I've spent this year tracking.

Sahil Bhaiwala, Chief Strategy and Innovation Officer at Handshake AI, published one of the sharpest breakdowns I've read this year on why the "data labeling" industry is misnamed and massively undervalued. Knowledge work has no chess-like win condition, so someone has to write the rules: design the scenario, build the environment, author the rubric that can grade any valid answer. He calls it data authoring, and he makes the case that it's a $100B+ market within three years, concentrated among a handful of players who can do the hard, capital-intensive parts of it at scale.

It's a rigorous piece, and rigorous is the right word. He scopes the argument carefully: he explicitly sets physical intelligence aside from the analysis, flags that robots will need "three-dimensional environment simulations and scenarios authored by experts," and says he'll take it on in a future post.

That's the discipline that makes the rest of the thesis worth trusting. It's also the most interesting open question in the piece.

Here's why the two problems don't collapse into one.

Everything in Bhaiwala's framework, the shift from labeling to authoring, the rubric as the definition of winning, the idea that the scarce input is expert judgment, not human hours, works when the task lives inside a simulated environment you can fully specify. A tool the agent can call. A database it can query. A verifier that reads the output and scores it. That's true when the "environment" is a replica of QuickBooks or a Bloomberg terminal, because software is finite. Every state it can be in is, in principle, enumerable.

A robot's environment is the physical world. It isn't finite, and it doesn't hold still. Friction changes with humidity. A part is a few millimeters off-spec from the CAD file. A human walks into the workspace. Authoring every physical state in advance the way you can author every financial-modeling task in advance isn't a smaller version of the same problem. It's a different problem, which is likely exactly why he set it aside rather than force a premature answer into an already-ambitious piece.

This is the split I've been tracking all year across the Physical AI work.

Software agents run on a cognition layer that gets smarter every time someone feeds it better training data, better rubrics, better evals. That's the layer Handshake, Scale, and Mercor are fighting over, and Bhaiwala's case for why value concentrates there is convincing. But a robot also has a control layer underneath it, the part making sensor-fusion and motor decisions in real time, with hard latency guarantees software agents never have to meet. Better authored data makes the cognition layer smarter. It doesn't make the control layer deterministic. Those are two different infrastructure problems, and the industry is only now starting to treat them as separate questions instead of one.

I'm actively engaged with a company building that second layer, the deterministic control infrastructure that has to exist underneath any amount of authored training data before a robot can act on it reliably. It's not a data problem in Bhaiwala's sense. It's a systems architecture problem, and it's the natural next chapter to his own thesis, not a challenge to it.

Both economies are real, and they're about to need each other.

The data-authoring economy Bhaiwala describes is underpriced and getting bigger fast. Sitting right next to it is a second one: the real-time infrastructure that lets a physical system act on any of that authored data safely. Neither replaces the other. The companies and investors who understand both layers exist, and treat them as complementary rather than substitutes, are the ones positioned well for what comes next. I'd genuinely like to see Bhaiwala's future post on this.

If you've thought about the physical-intelligence data problem: does it get solved with better simulation, cheaper real-world collection, or something structurally different from what's working for knowledge work?

Reach me at arif@faris-capital.com

This perspective was originally published on LinkedIn. Join the discussion, add your thoughts, and follow for regular updates.

View on LinkedIn