On-device was a cost decision before it was a privacy one

We did not put models on the laptop because it sounded principled. We did it because the arithmetic of an Indian subscription forced it.

Redrob · July 28, 2026 · 5 min read

On-device inference has a reputation as a principled choice. Data stays with the user, nothing goes to a server, privacy by architecture rather than by promise.

All of that is true of Redrob Office and none of it is why we built it that way.

We built it that way because of arithmetic.

The line item nobody removes

Every agent that answers from the cloud carries a token bill on every query. It is a variable cost, it scales with success, and it never goes away. The better the product does, the larger it gets.

For most software companies that is survivable, because the price they charge has room in it. For a product priced at what the Indian market will actually pay, it is not room, it is the whole ceiling. Serving cost and price sit close enough together that the gap between them is the business, and anything that widens the gap is worth more than a feature.

Moving inference onto the device does not reduce that line item. It deletes it.

The user's laptop is already bought and already powered. Its compute is a sunk cost belonging to somebody else, and the marginal cost to us of using it is zero. Not lower. Zero. There is no other lever in the stack that does that.

Why documents are the right place to try it

On-device only works when the model can be small, and the model can only be small when the task is narrow enough not to need a large one.

Office is documents, code and workflow. That work has a property that general chat does not: the context is right there. The document is on the machine. The repository is on the machine. The question is almost always about something local, and the model is not being asked to know things about the world, it is being asked to work with material it can already see.

Retrieval does the heavy lifting, and the model does the reasoning over what retrieval found. That combination fits in a much smaller model than a system expected to answer from memory, and small enough is the entire requirement.

General conversation is the opposite case. It is unbounded, it draws on everything, and it does not fit. Which is why Office runs locally and Chat does not, and why we did not pretend the same architecture suited both.

What we gave up

Three things, and they are real.

Model size is capped by the worst machine we support. The ceiling is not set by the best laptop; it is set by the ordinary one. Every capability decision runs into that wall.

Updates are a distribution problem again. A cloud model improves for everybody the moment we deploy it. A local model improves when the user takes the update, which means we are back in a world of versions and long tails.

Debugging happens somewhere we cannot see. When a cloud answer is wrong we have the trace. When a local answer is wrong we have a report from a person describing what they remember.

We took all three, because none of them is as expensive as the line item.

The consequence we did not plan for

Then the Digital Personal Data Protection Act arrived and made the same choice look prescient.

If inference happens on the device, the document never leaves it. There is no cross-border transfer to justify, no processing agreement to negotiate, no data residency question to answer, because the question does not arise. Compliance is a property of the architecture rather than a set of controls layered on top of it.

That is genuinely valuable, and it is the reason enterprise conversations about Office go faster than they used to.

But it was a consequence, not a motive, and we would rather say so. A cost decision that turns out to have a privacy dividend is a good outcome. Describing it afterwards as a privacy decision would be a nicer story and a less useful one, because the actual lesson generalises: the constraint that forces the unusual architecture is usually the one worth listening to.

BACKED BY

Korea Investment Partners

KB Investment

Kiwoom Investment

KDB Capital

DS&Partners

Murex Partners

Daekyo Investment

Wanted Lab

© 2026 Redrob. All rights reserved.

Privacy

Terms

Security

English