Newsroom — AI
AI for computer use
AI agents can now operate a computer the way a person does — seeing the screen, moving the mouse, typing. What that changes, where it breaks, and why the agent belongs inside the work, not on top of the operating system.
By VewMet — October 2026 — 5 min read
Contents
The interface flips
For forty years, software assumed a human at the controls: you learn the app, you click the buttons. The new wave — Anthropic's computer use, OpenAI's Operator, and a fast-moving field of agent startups — inverts that. The model learns the app. It looks at the screen the way you do, moves the pointer, types into fields, and works through the task.
This is a genuine break from the last decade of automation. Instead of one integration per app — brittle, maintained by someone, always slightly behind — you get one general skill: see the interface, use the interface. The long tail of software, the part with no API and no integration budget, suddenly becomes operable.
Consider the expense report. Nobody's finance team is going to build an integration for the rickety travel portal the company adopted in 2019. But an agent that can open the portal, read the form the way an employee would, and file the claim correctly has just automated something no API ever reached. Multiply that by every internal tool, government portal, and legacy dashboard in existence, and the scale of the unlock becomes clear.
The loop, in miniature
Strip away the demos and every computer-use agent is this loop:
# The computer-use loop, in miniature. Illustrative, not production code.
while not task.done and steps < MAX_STEPS:
screenshot = desktop.capture() # see the interface
action = model.decide(screenshot, task) # click, type, scroll, or done
if action.risk == "consequential": # no undo for the big ones
human.confirm(action) # the checkpoint demos skip
desktop.execute(action)
steps += 1Screenshot, decide, act, repeat — with a human checkpoint before anything irreversible. Everything the rest of this piece argues about lives inside those eleven lines: the generality, the slowness, and the trust problem.
Why it matters for knowledge work
Real knowledge work doesn't live inside one app. It lives in the seams: the spreadsheet that feeds the slide deck, the inbox that feeds the doc, the dashboard whose numbers get retyped into a report. That seam-work is copy-paste labor — high attention, zero leverage.
A computer-using agent attacks exactly that. It doesn't need the seams to be engineered away first. It can do the unglamorous traversal today: gather, transcribe, reconcile, file.
That is worth appreciating. The boring middle of office work has resisted automation for decades because it was never worth integrating. Now it might not need integrating at all.
There is a second-order effect worth naming. When the cost of operating software drops toward zero, the software itself stops being the moat. The moat moves to the work: the judgment, the relationships, the taste. Tools that assumed captive users will have to compete on the experience instead of the lock-in. That is a healthy trade for everyone except the vendors who were coasting.
Where it breaks
And yet: a 95%-reliable agent is not a tool, it's a liability. Desktop control has no undo for the consequential things — the wrong file deleted, the wrong message sent, the wrong form submitted. The industry's favorite demo is a tidy task; the real world is pop-ups, auth walls, CAPTCHAs, and interfaces that changed overnight.
Then there's the cost of watching paint dry. These agents are slow — they reason, screenshot, click, wait, re-reason — and every step burns tokens. A task you do in ninety seconds can take an agent twenty minutes and cost real money. The economics only work for tasks where human attention is genuinely more expensive than the agent's wandering, which is a smaller set than the demos suggest.
And the deepest problem is trust. An agent that can move your mouse can move your money. The permission model for "do what I would do" doesn't exist yet; we're granting it anyway, one demo at a time. Every new capability arrives before the vocabulary to constrain it. Sandboxing, scoped credentials, human-in-the-loop checkpoints — the industry is building the brakes while the car is already on the highway.
The industry wants autopilot for your desktop. We think that's the wrong target.
Our take: put the agent where the work is
Here's the builder's view from VewMet. Computer use is real, and it will matter — but the winning shape isn't a floating agent with the keys to your whole machine. It's an agent that lives inside the surface where the work already happens.
For writing and collaboration, that surface is the document. inkk.ai is built on this premise: real-time co-editing, a side-by-side chat copilot, audio and video calling, and a marketplace of agents — all in the same line you're writing. The agent doesn't need your desktop. It needs your context: the paragraph, the draft, the collaborators, the thread.
A document is a natural sandbox. Actions are scoped, visible, and reversible. Your collaborators can see what the agent changed, down to the letter. Compare that to an agent clicking around your OS while you watch a screenshare of your own machine and hope.
There is also a quieter advantage: accountability. When an agent acts inside a shared document, its work is reviewable by the team in the same way a colleague's draft is. Suggestions, edits, and agent actions all live in one history. The desktop agent, by contrast, leaves its trail in a log file nobody reads.
The pitch is that agents will replace knowledge workers. We disagree. The good version of this technology doesn't remove the human from the loop — it removes the seams between the human's tools.
What we're watching
We'll be watching reliability benchmarks, not demos. The day a computer-using agent can do a messy, interrupted, real-world task ten times in a row without supervision is the day the conversation changes. We want to see numbers on recovery from unexpected states — the pop-up, the changed layout, the expired session — because that is where the demos end and the product begins.
We're also watching the permission story. Whoever ships a legible, enforceable model for "the agent may do X but must ask before Y" will unlock the enterprise half of this market overnight. Until then, the technology stays in the demo tier for anything consequential.
Until both arrive, we're building where we are: the collaboration surface. Agents inside the document, beside the writer, accountable to the team. The desktop can wait.
Keep reading
Engineering — October 2026
Jev: the System One model
TypeSafe AI's Jev doesn't generate text. It answers typed questions with calibrated probabilities — fast, cheap judgment for agent stacks. Why the missing layer in AI products isn't a bigger model, it's a faster decider.
Product — February 2025
inkk.ai Marketplace is live
Our AI agent marketplace launched in February 2025 — AI agents that take care of all your writing needs.
Product — October 2026
OpenAI just announced Pages. We've been shipping it since February.
OpenAI's Pages is a collaborative AI document editor. Ours has been live since February — with caret-level attribution, in-document calls, and zero vendor lock-in.
Product
VewMet Calls
Imagine AI joining your calls as a participant — listening, moderating, and contributing in near real-time. Launching soon.