Skip to main content

Demo: Computer Use

V2 and above. Part of the invite-only hosted editions, not the free self-hosted V1. See the editions.

What you'll see​

You give AI Partner a task that requires interacting with a website. It picks the right browser tier for the job — fast DOM extraction for standard pages, vision-based navigation for complex SPAs, a full containerized desktop for the hardest cases — and completes the task autonomously.


The four tiers​

TierMethodSpeedUsed for
T1Reads the page directly, no screenshot needed~2sOrdinary pages with readable content
T2Looks at a screenshot of the page and works out where to click~10sPages built entirely in the browser, where there's nothing to read
T3A full cloud desktop with a real browser~30sJoining meetings, CAPTCHA-protected flows, anything needing a whole desktop
T4The real keyboard and mouse of the host machine~2s/actionNative desktop apps and real signed-in sessions (restricted to named apps, single-user only)

The agent always starts at T1 and escalates automatically if a lower tier fails.

T4 is the only tier that drives an actual desktop. Two separate gates must both be open before it runs; dragging the cursor into a screen corner aborts it; it backs off the moment you touch the mouse; and every action is written to a durable audit trail. Read the full Host Desktop Control guide before enabling it — and on a hosted instance, use Your Own Computer instead.


The cloud desktop remembers its own setup​

When a cloud desktop session is active, commands run inside that same desktop — the one holding the browser you're watching — rather than in a throwaway environment that's discarded between steps.

Why this matters:

BeforeAfter
Something installed in step 1 is gone by step 2 (a fresh container each time)Step 2 still has it — same session throughout
Changing directory didn't carry to the next commandThe working directory persists across actions
Environment settings were lost between commandsThey survive for the whole session
Restarting the session lost everything installedYour session's own storage persists across restarts

How it works:

  • The partner runs in a full, isolated Linux environment with a persistent shell — it can install what it needs and reach the internet, and what it installs is still there on the next step.
  • It keeps your working directory and environment between commands, so multi-step work flows naturally.
  • When a partner session is active, the agent runs commands inside that environment automatically — no special syntax needed.

This is the single biggest unlock for specialist work. The agent installs whatever a task needs, and it stays installed for as long as the session is up. One general cloud desktop can therefore become a security-testing workstation, a data-analysis notebook or a build environment — decided by the job, not fixed in advance.


Demo 1: Standard web extraction (T1)​

Type this:

Go to https://news.ycombinator.com and extract the top 10 stories.
For each story get: title, URL, points, and comment count.
Return them as a clean numbered list.

What happens:

✅ T1: browser_navigate(https://news.ycombinator.com)
✅ T1: browser_extract(selector: ".athing, .subtext")
→ Extracted 10 stories in 1.8 seconds
✅ Formatted and returned

T1 reads the DOM directly — no screenshot, no LLM vision — so it's extremely fast.


Demo 2: Complex SPA navigation (T2)​

Type this:

Go to https://linear.app and navigate to the pricing page.
Extract all plan names, prices, and the features listed under each plan.

What happens:

Linear's pricing page is built entirely in the browser — there's no readable page to extract. T1 comes back with nothing useful, so the agent escalates:

⚠️ T1: extraction returned empty data → escalating to T2
✅ T2: browser_navigate(https://linear.app/pricing) with stealth mode
✅ T2: browser_screenshot() → captured page
✅ T2: vision_analyze(screenshot) → "I can see 3 pricing tiers: Free, Business, Enterprise..."
✅ T2: browser_extract(targeted selectors based on visual analysis)
→ Extracted 3 plans, prices, and feature lists
✅ Formatted and returned

T2 takes a screenshot, uses the LLM's vision to understand the page layout, then extracts the data using the selectors it identifies visually.


Demo 3: CAPTCHA handling (T2 → user handoff)​

Type this:

Go to https://www.linkedin.com/in/satya-nadella and extract his current job title,
company, location, and latest 3 posts.

LinkedIn aggressively blocks automated browsers. When the agent encounters a CAPTCHA or login wall:

✅ T2: browser_navigate(https://linkedin.com/in/satya-nadella)
⚠️ T2: CAPTCHA detected — pausing for human handoff

In the AI Partner UI, you'll see:

  • A live screenshot of the CAPTCHA page
  • A "Take Control" button

Click Take Control → a browser window opens on your machine → solve the CAPTCHA → click Continue in AI Partner.

✅ Resumed after CAPTCHA solved by user
✅ T2: browser_extract(profile data)
→ Title: CEO, Microsoft
→ Location: Redmond, WA
→ Latest posts extracted

The agent re-validates the current page state after the handoff before extracting data — it confirms it's actually on the profile page, not a post-CAPTCHA redirect.


Demo 4: Form filling (T1)​

Type this:

Go to https://formspree.io/forms/new and fill in the form to create a new form endpoint.
Use these values:
- Form name: Test Form
- Email: test@example.com
Submit the form and tell me the endpoint URL that appears after submission.

What happens:

✅ T1: browser_navigate(https://formspree.io/forms/new)
✅ T1: browser_fill(selector: "#name", value: "Test Form")
✅ T1: browser_fill(selector: "#email", value: "test@example.com")
✅ T1: browser_click(selector: "button[type=submit]")
✅ T1: browser_extract(selector: ".endpoint-url")
→ Endpoint: https://formspree.io/f/xyzabc
✅ Result returned

Demo 5: Meeting join (T3 container)​

Meeting attendance is the most demanding case. The cloud desktop starts a full screen-and-sound environment, joins the meeting in a real browser, and captures the audio.

See the full walkthrough: Meeting Attendance demo →


Configuring computer use tiers​

Go to Settings → Computer Use to configure:

SettingDefaultNotes
Default tierAuto-escalateStart at T1, escalate on failure
T4 host controlDisabledEnable only for single-user, trusted environments — refused entirely when auth is on. See T4 guide
T4 max steps20Hard cap to prevent runaway host control
T4 allowed applicationsA short default listThe focused window must match one of them before any action runs (empty = any app)
T4 denylistpassword managersAlways blocked, even with an open allowlist
T4 approval modesensitivenever / sensitive / always — pause for approval before risky host actions
T4 media keysenabledmedia_key verb: volume, mute, play/pause, next/previous (bypasses allowlist)
T4 drift threshold50 pxT4 pauses if you move your cursor between its actions
T4 monitor captureprimaryOn multi-monitor systems, only the primary monitor is captured
T4 live broadcastOff (live_stream_fps: 0)Set to 5-10 fps to see a continuous video of your screen in the Inspector while T4 runs
CAPTCHA timeout5 minutesHow long to wait for user to solve CAPTCHA
Stealth modeOn for T2Rotates user-agent and headers to avoid detection

T4 emergency controls​

ControlHow to trigger
Hard abort (FAILSAFE)Drag your mouse cursor to any screen corner
Kill-now (all tiers)The STOP button. Stops the current work instantly — at every tier, not just T4 — and also fires if you cancel the parent goal. No timeout involved: it stops when you press it
Inspect audit logThe durable record of every action taken, with what was passed to each one
Drift guardAutomatic — move the cursor mid-task and T4 pauses for you

What computer use can't do​

Computer use is not magic. Some limitations:

  • Two-factor authentication: if a site requires 2FA, the agent pauses for HITL (you enter the code)
  • PDF downloads in containers: downloaded files are extracted from the container and saved to workspace
  • Paid sites: the agent can't pay for access it doesn't have
  • Sites with Cloudflare turnstile: T2 may solve simple CAPTCHAs, complex ones always escalate to you
  • Desktop apps (non-web): T4 only, and only for allowlisted applications

Combining with goal execution​

Computer use is one tool among many in the ReAct loop. A single goal can mix browser, Python, and file generation:

Scrape the Stripe pricing page and extract all plan details.
Then scrape the Paddle pricing page and do the same.
Compare the two and generate a side-by-side Excel spreadsheet.

The agent handles this as:

  1. T1: scrape Stripe pricing → Python: structure data
  2. T1: scrape Paddle pricing → Python: structure data
  3. generate_excel: create comparison table
  4. Files panel: download link available