Host Desktop Control
V2 and above. Part of the invite-only hosted editions, not the free self-hosted V1. See the editions.
What this is, and what it isn't
Host control — T4 — drives the real mouse and keyboard of the machine AI Partner itself runs on. Every other way of using a computer is sandboxed: a browser context, or an isolated container. If those go wrong, you close a tab or restart a container.
T4 has no such escape hatch. Every click lands on that actual desktop, and every keystroke goes into whichever window has focus.
In return, it's the only option that drives native applications — Office, a desktop Slack, an IDE, a file manager, anything installed locally — using the real signed-in sessions already on that machine.
T4 is single-user and opt-in. Never enable it on a machine you share with anyone else.
As a hard guarantee it is refused outright on any instance with sign-in enabled — a multi-user deployment can never drive the host desktop, whatever the configuration says. Users of a shared instance get an isolated cloud desktop instead, or their own computer via a connected companion.
Working on a hosted instance? T4 is not your route — Your Own Computer is. That connects your laptop to a hosted AI Partner, with per-folder permission and a banner on the machine while it works. T4 is for the case where AI Partner is running on the very machine you want it to drive.
Choosing between a cloud desktop and host control
| Cloud desktop | Host control (T4) | |
|---|---|---|
| Isolation | Complete — its own container | None — the real machine |
| A mistake is recoverable | Yes, discard the container | No — real keystrokes, really typed |
| Drives native desktop apps | No | Yes |
| Uses your real signed-in sessions | No | Yes |
| Live view | Yes, by default | Optional, off by default |
| Suitable for multi-user | Yes | No |
| Time to first action | ~20s | ~2s |
Default to the cloud desktop. Reach for host control only when the task genuinely needs a native application or the real session state on that machine.
The constraint everything follows from
You and T4 share one mouse and one keyboard. You cannot both drive that machine at once. That's physics, not a design decision, and every behaviour below exists because of it.
Turning T4 on does not mean it's running. It means the system is allowed to use it if a specific goal asks for it. It takes your input devices only during an active attempt.
| State | Who has the mouse | What you experience |
|---|---|---|
| Enabled, nothing running | You | Normal — T4 might as well not exist |
| Enabled, attempt running | T4 — step back | The cursor moves on its own, windows get clicked, text gets typed |
| Enabled, attempt paused | You | It's waiting for you |
Two ways it gets deployed
How you supervise depends entirely on which of these you're in. Decide before you configure the live view.
AI Partner runs on your own machine, and its window is open on the same desktop it's driving.
- You see everything directly — its mouse is your mouse, moving.
- The live view is just a mirror of the screen you're already looking at.
- Put the AI Partner window on a second monitor, or shrink it into a corner, so it isn't covering the apps being driven.
- Turn the live view off. There's nothing to broadcast — the live screen is your screen.
Your controls
Layered deliberately: different mechanisms fail in different ways, so there's more than one.
| Control | What it does | Best for |
|---|---|---|
| Inspector | Live view of the screen being driven | Watching it work |
| Activity log | A step-by-step written trace | Reading afterwards |
| STOP | Kills the current attempt instantly — not at the end of the step. Works at every tier, and fires automatically if you cancel the parent goal | "That's wrong", while the app is reachable |
| Audit trail | A durable record of every action, with timestamps and what was passed | Post-mortem and compliance |
| FAILSAFE corner | Drag the cursor into any screen corner and the next action aborts. Physical, needs no app, no network and no working software | The panic button when the app is frozen or the network is down |
| Drift guard | Touch the mouse and it backs off by itself | "Let me just check something" |
| It asks for help | It pauses on its own for a sign-in, a CAPTCHA or anything needing your judgement | When it hits a wall |
The FAILSAFE corner
Learn this one reflex: if it's doing something wrong, throw your mouse into a corner. Nothing to click, nothing to find, no dependency on any software still working. It's the control that survives everything else failing.
The drift guard
It notes where it left your cursor. Before the next action it looks again — and if you've moved it, it pauses and waits for you to finish, rather than fighting you for the pointer. If you don't come back, it gives up rather than resuming behind you.
Grab your mouse mid-task and it backs off. That's the whole behaviour.
A drifting touchpad or a jittery wireless mouse can trip the guard constantly. Raise the sensitivity threshold in settings rather than turning the guard off.
Restricting what it can touch
Allowed applications
You can restrict T4 to acting only when the focused window is one of a list of applications you name.
This matters more than it sounds. With a list set, a popup that steals focus mid-task causes T4 to stop — rather than typing your password into whatever dialog just appeared. Leave the list empty and it may act in any window.
Every blocked attempt is recorded in the audit trail, along with which window was actually in focus — which is also how you correct a list entry that isn't matching.
Always-blocked applications
The inverse, and it wins over everything else: window patterns that are refused even when the allowed list is wide open. It ships defaulting to the common password managers, so the agent can never drive your vault.
Asking before sensitive actions
| Mode | Behaviour |
|---|---|
| Never | No approvals. The allowed and blocked lists still apply. |
| Sensitive (default) | Asks before risky actions and in sensitive windows — shell commands that destroy things, run dialogs, terminals, banking and payment windows |
| Always | Asks before every action. Maximum oversight, slowest going. |
When an action needs approval and nobody is connected to approve it, it is blocked and reported — never quietly allowed. That's the direction a fail-safe has to point.
What the prompt looks like
The approval bar docks into the Inspector, over the live screen you're already watching, with a pulsing marker on the exact spot it intends to click — labelled with which monitor, if you're capturing more than one. You get Approve / Deny / Stop and a countdown.
┌─ Inspector ───────────────────────────────┐
│ [ live screen ] ◎ ← target marker │
├────────────────────────────────────────────┤
│ ⚠ type "sudo rm -rf" → terminal ⏱0:54 │
│ [ ✓ Approve ] [ ✗ Deny ] [ ⛔ Stop ] │
└───────────────────────────────────── ───────┘
If the countdown runs out it denies, not approves. An unattended prompt never becomes consent.
Volume and playback keys are exempt from the lists and the approval gate — they're global and harmless. They still obey Stop and are still written to the audit trail.
The live view
Off by default. Turned on, you get a continuous video feed of the machine's screen in the Inspector — the same panel where browser and cloud-desktop frames already appear, with a LIVE indicator while it's streaming.
Why have a video feed when the agent takes its own screenshots? Because they're for different people:
- The agent's screenshots are taken at decision points — one frame per step, which is all it needs.
- The live view is continuous, and it's for you. Animations, transitions, focus changes, a dialog appearing and vanishing — visible to you, not necessarily to the agent.
Live view plus the FAILSAFE corner plus the drift guard is what makes this feel like supervising a colleague rather than launching a script and hoping.
Multiple monitors
By default it captures and acts on the primary monitor only. The vision model gets one coherent screenshot, and clicks land where it expects them to.
You can switch to capturing the whole desktop across every monitor, but only do it if a task genuinely has to span them: vision models get worse at very wide images, so three monitors side by side is a harder problem than one, not an easier one.
What it can do
The vision model chooses from a full desktop vocabulary: click (single, double, triple, right, middle), type text, press keys and chords, hold a key, scroll, hover, drag, query the cursor, wait, and the media keys. Plus two it uses to end an attempt honestly — ask for human help, and report failure — because a run that couldn't finish should say so rather than claim otherwise.
Turning it on
T4 is off by default, and two separate gates must both be open before it will run.
- 1Install what the machine needs
Host control needs a small set of desktop-automation components on that machine. On macOS you'll also need to grant it Accessibility permission — macOS will not let anything drive the desktop without it.
- 2Enable it in your instance settings
Turn host control on, then set what it's allowed to do: which applications are permitted, which are always refused, when to ask you, drift sensitivity, monitor capture and the live view.
- 3Opt in per goal
Being enabled isn't enough. Each goal that should reach the host desktop has to ask for it specifically. A goal that didn't ask can never end up driving your machine by accident.
- 4Learn the corner
Before you run anything real: the emergency stop is dragging your cursor to a screen corner. Practise it once.
Troubleshooting
FAILSAFE fires immediately on the first action
Your cursor is already sitting in a corner. Move it to the middle of the screen and try again.
It pauses constantly and never gets anywhere
The drift guard is being tripped by a touchpad or a jittery wireless mouse. Raise the drift threshold in settings.
Everything is blocked
Your allowed-application entries don't match the real window names. Check the audit trail — blocked entries record the window that was actually focused. Use part of that name as your entry.
No live view in the Inspector
Confirm the live view is switched on for that machine, that an attempt is genuinely running, and that the Inspector panel is open. The LIVE indicator tells you whether anything is actually being broadcast.
It says host control isn't available at all
Expected on any instance with sign-in enabled — that's the hard guarantee, not a misconfiguration. Use a cloud desktop, or connect your own computer.