Demo: Computer Use
V2 and above. Part of the invite-only hosted editions, not the free self-hosted V1. See the editions.
What you'll see
You give AI Partner a task that requires interacting with a website. It picks the right browser tier for the job — fast DOM extraction for standard pages, vision-based navigation for complex SPAs, a full containerized desktop for the hardest cases — and completes the task autonomously.
The four tiers
| Tier | Method | Speed | Used for |
|---|---|---|---|
| T1 | Reads the page directly, no screenshot needed | ~2s | Ordinary pages with readable content |
| T2 | Looks at a screenshot of the page and works out where to click | ~10s | Pages built entirely in the browser, where there's nothing to read |
| T3 | A full cloud desktop with a real browser | ~30s | Joining meetings, CAPTCHA-protected flows, anything needing a whole desktop |
| T4 | The real keyboard and mouse of the host machine | ~2s/action | Native desktop apps and real signed-in sessions (restricted to named apps, single-user only) |
The agent always starts at T1 and escalates automatically if a lower tier fails.
T4 is the only tier that drives an actual desktop. Two separate gates must both be open before it runs; dragging the cursor into a screen corner aborts it; it backs off the moment you touch the mouse; and every action is written to a durable audit trail. Read the full Host Desktop Control guide before enabling it — and on a hosted instance, use Your Own Computer instead.
The cloud desktop remembers its own setup
When a cloud desktop session is active, commands run inside that same desktop — the one holding the browser you're watching — rather than in a throwaway environment that's discarded between steps.
Why this matters:
| Before | After |
|---|---|
| Something installed in step 1 is gone by step 2 (a fresh container each time) | Step 2 still has it — same session throughout |
| Changing directory didn't carry to the next command | The working directory persists across actions |
| Environment settings were lost between commands | They survive for the whole session |
| Restarting the session lost everything installed | Your session's own storage persists across restarts |
How it works:
- The partner runs in a full, isolated Linux environment with a persistent shell — it can install what it needs and reach the internet, and what it installs is still there on the next step.
- It keeps your working directory and environment between commands, so multi-step work flows naturally.
- When a partner session is active, the agent runs commands inside that environment automatically — no special syntax needed.
This is the single biggest unlock for specialist work. The agent installs whatever a task needs, and it stays installed for as long as the session is up. One general cloud desktop can therefore become a security-testing workstation, a data-analysis notebook or a build environment — decided by the job, not fixed in advance.
Demo 1: Standard web extraction (T1)
Type this:
Go to https://news.ycombinator.com and extract the top 10 stories.
For each story get: title, URL, points, and comment count.
Return them as a clean numbered list.
What happens:
✅ T1: browser_navigate(https://news.ycombinator.com)
✅ T1: browser_extract(selector: ".athing, .subtext")
→ Extracted 10 stories in 1.8 seconds
✅ Formatted and returned
T1 reads the DOM directly — no screenshot, no LLM vision — so it's extremely fast.
Demo 2: Complex SPA navigation (T2)
Type this:
Go to https://linear.app and navigate to the pricing page.
Extract all plan names, prices, and the features listed under each plan.
What happens:
Linear's pricing page is built entirely in the browser — there's no readable page to extract. T1 comes back with nothing useful, so the agent escalates:
⚠️ T1: extraction returned empty data → escalating to T2
✅ T2: browser_navigate(https://linear.app/pricing) with stealth mode
✅ T2: browser_screenshot() → captured page
✅ T2: vision_analyze(screenshot) → "I can see 3 pricing tiers: Free, Business, Enterprise..."
✅ T2: browser_extract(targeted selectors based on visual analysis)
→ Extracted 3 plans, prices, and feature lists
✅ Formatted and returned
T2 takes a screenshot, uses the LLM's vision to understand the page layout, then extracts the data using the selectors it identifies visually.
Demo 3: CAPTCHA handling (T2 → user handoff)
Type this:
Go to https://www.linkedin.com/in/satya-nadella and extract his current job title,
company, location, and latest 3 posts.
LinkedIn aggressively blocks automated browsers. When the agent encounters a CAPTCHA or login wall:
✅ T2: browser_navigate(https://linkedin.com/in/satya-nadella)
⚠️ T2: CAPTCHA detected — pausing for human handoff
In the AI Partner UI, you'll see:
- A live screenshot of the CAPTCHA page
- A "Take Control" button
Click Take Control → a browser window opens on your machine → solve the CAPTCHA → click Continue in AI Partner.
✅ Resumed after CAPTCHA solved by user
✅ T2: browser_extract(profile data)
→ Title: CEO, Microsoft
→ Location: Redmond, WA
→ Latest posts extracted
The agent re-validates the current page state after the handoff before extracting data — it confirms it's actually on the profile page, not a post-CAPTCHA redirect.
Demo 4: Form filling (T1)
Type this:
Go to https://formspree.io/forms/new and fill in the form to create a new form endpoint.
Use these values:
- Form name: Test Form
- Email: test@example.com
Submit the form and tell me the endpoint URL that appears after submission.
What happens:
✅ T1: browser_navigate(https://formspree.io/forms/new)
✅ T1: browser_fill(selector: "#name", value: "Test Form")
✅ T1: browser_fill(selector: "#email", value: "test@example.com")
✅ T1: browser_click(selector: "button[type=submit]")
✅ T1: browser_extract(selector: ".endpoint-url")
→ Endpoint: https://formspree.io/f/xyzabc
✅ Result returned
Demo 5: Meeting join (T3 container)
Meeting attendance is the most demanding case. The cloud desktop starts a full screen-and-sound environment, joins the meeting in a real browser, and captures the audio.
See the full walkthrough: Meeting Attendance demo →
Configuring computer use tiers
Go to Settings → Computer Use to configure:
| Setting | Default | Notes |
|---|---|---|
| Default tier | Auto-escalate | Start at T1, escalate on failure |
| T4 host control | Disabled | Enable only for single-user, trusted environments — refused entirely when auth is on. See T4 guide |
| T4 max steps | 20 | Hard cap to prevent runaway host control |
| T4 allowed applications | A short default list | The focused window must match one of them before any action runs (empty = any app) |
| T4 denylist | password managers | Always blocked, even with an open allowlist |
| T4 approval mode | sensitive | never / sensitive / always — pause for approval before risky host actions |
| T4 media keys | enabled | media_key verb: volume, mute, play/pause, next/previous (bypasses allowlist) |
| T4 drift threshold | 50 px | T4 pauses if you move your cursor between its actions |
| T4 monitor capture | primary | On multi-monitor systems, only the primary monitor is captured |
| T4 live broadcast | Off (live_stream_fps: 0) | Set to 5-10 fps to see a continuous video of your screen in the Inspector while T4 runs |
| CAPTCHA timeout | 5 minutes | How long to wait for user to solve CAPTCHA |
| Stealth mode | On for T2 | Rotates user-agent and headers to avoid detection |
T4 emergency controls
| Control | How to trigger |
|---|---|
| Hard abort (FAILSAFE) | Drag your mouse cursor to any screen corner |
| Kill-now (all tiers) | The STOP button. Stops the current work instantly — at every tier, not just T4 — and also fires if you cancel the parent goal. No timeout involved: it stops when you press it |
| Inspect audit log | The durable record of every action taken, with what was passed to each one |
| Drift guard | Automatic — move the cursor mid-task and T4 pauses for you |
What computer use can't do
Computer use is not magic. Some limitations:
- Two-factor authentication: if a site requires 2FA, the agent pauses for HITL (you enter the code)
- PDF downloads in containers: downloaded files are extracted from the container and saved to workspace
- Paid sites: the agent can't pay for access it doesn't have
- Sites with Cloudflare turnstile: T2 may solve simple CAPTCHAs, complex ones always escalate to you
- Desktop apps (non-web): T4 only, and only for allowlisted applications
Combining with goal execution
Computer use is one tool among many in the ReAct loop. A single goal can mix browser, Python, and file generation:
Scrape the Stripe pricing page and extract all plan details.
Then scrape the Paddle pricing page and do the same.
Compare the two and generate a side-by-side Excel spreadsheet.
The agent handles this as:
- T1: scrape Stripe pricing → Python: structure data
- T1: scrape Paddle pricing → Python: structure data
- generate_excel: create comparison table
- Files panel: download link available