V3 turns the agent into a twin: it remembers you permanently, disagrees with dated evidence, challenges you when you drift from your own goals โ and earns, approval by approval, the right to act in your name. You grow. It grows. Both, measurably.
Every other agent resets to zero each session. A twin compounds โ and it compounds you with it.
% of delegated work completed without a human touch ยท autonomy is earned, never granted โ one edit revokes it
Not a persona file that sounds consistent and knows nothing. Structure โ because for a twin, care is structure or it's absent.
CHARTER.md โ a constitution you own, in your workspace. You grant which domains it may confront you about, evidence thresholds, and back-off terms. Deny by default; amended by ceremony; hot-reloaded on every action.
The founding episode, turning points, accepted corrections, granted authorities โ permanent, never summarized away, present in every prompt. Carried state always arrives with its date: no date, nothing carried.
It holds evidence-backed claims about you and pushes back once when you contradict yourself โ citing dates. Overruled? It complies and keeps the question open. Corrected? Revised out loud, prior claim kept in the audit trail.
At most one challenge a week, only through a deterministic gate: granted domain, evidence threshold met. Back off and it obeys AND remembers; the third back-off retires the goal forever. Timeout records nothing โ silence is never consent.
Clean-approval streaks graduate action classes to act-without-asking โ only within charter scope, only when you're unreachable (by default), always receipted. One edit revokes everything and the streak restarts from zero.
Everything done in your absence, each action with a receipt, delivered at 09:00. Nothing happened? No message. A twin is a partner, not a notification app.
One view in the app: what needs you (answer with the same buttons as Telegram), the relationship's full activity record, open threads, positions with evidence, live trust streaks, and your charter.
In team deployments every member gets their own isolated twin โ and admins see what twins did (actions, receipts, freeze switches), never what they know. Memories, reasons, and goals stay owner-private, structurally.
You choose the accountability domains in a charter you own. The twin shows โ per domain โ how well it understands the pattern, how close it is to being cleared to act, whether its nudges land, and which way each trends. Board-ready, not a black box.
| Domain (you choose these) | Understands | Cleared to act | Nudges landing | 90d |
|---|---|---|---|---|
Goals vs. actual activity stated_goals_vs_activity |
4 | 2/2 | โฒ | |
Delivery vs. plan project_build_arcs |
3 | 1/1 | โฒ | |
Commitments to clients & team commitments_to_others |
2 | 1/2 | โฒ | |
Pipeline follow-through pipeline_followthrough |
1 | 0/0 | ยท | |
Cash & runway discipline runway_discipline |
2 | 1/1 | โฒ |
You set what it's allowed to hold people to โ a founder picks pipeline & runway; a delivery lead picks plan & client commitments โ and get an auditable read of whether the AI is genuinely learning the operation: positions revised, corrections that stick, autonomy earned and sustained, nudges that land.
V3 ships as the twin-connect update inside both editions โ V1 for yourself, V2 for your whole team.
| Capability | V1 โ Agent | V2 โ Team Workforce | V3 โ Digital Twin |
|---|---|---|---|
| Memory | Vector + RAG | 5-layer (episodic, biographic, counterparty) | โ + Constitutional โ permanent, dated, never compressed |
| Holds positions about you | โ | โ | โ evidence-backed, revisable, audited |
| Disagrees with you | โ | โ | โ once, citing dates; overruled โ forgotten |
| Challenges your drift | โ | โ | โ weekly, charter-gated, obey-and-remember |
| Acts in your name | โ | Approval-gated proxy (inbox, meetings, phone) | โ earned per action class; one edit revokes |
| Reports its absence work | โ | โ | โ daily digest with receipts |
| Org governance | โ | Admin console, audit, invites | โ conduct-not-content: actions visible, relationship private |
| Goal execution & validation | โ evidence-based | โ enhanced | โ same engine underneath |
Get first access to the twin update, onboarding for your charter, and direct founder contact.
Your twin only ever answers to you โ starting with this list.
AI Partner V1 is a self-hosted, open-source autonomous agent. Give it a goal in plain English โ it researches, codes, generates documents, and delivers the finished result on your own machine. Want the twin that grows with you? Meet V3 โ
Agent pauses before running any script. Approve or reject via Telegram inline buttons. Also handles CAPTCHA hand-off, clarifications, and mid-execution input requests.
Full browser with 15fps live screencast. SPA-aware rendering, stealth mode. CAPTCHA detected? Agent pauses, you solve in the live panel, it resumes.
Every script runs inside a persistent container with Python, Node.js, and Bash. pandas, yfinance, matplotlib pre-loaded. Network isolation per-goal.
PDF via Chromium, multi-sheet XLSX, PowerPoint, and DOCX โ all production-ready, tracked in the dashboard, and instantly downloadable.
After a successful goal, the system generalizes the solution into a reusable parameterized template. Versioned, deduplicated, and promotable to a callable tool.
Orchestrators spin up to 5 parallel sub-agents with budget controls. When a research agent exhausts its cap, it writes a handoff summary โ the parent picks up seamlessly.
Cron-expression tasks persist across restarts. The Proactive Agenda reads your task list and memory โ uses LLM judgment to pick the best action per tick.
Episodic log, vector search, persona traits, and periodic consolidation. Works with Ollama, OpenAI, Cohere, or pure-JS TF-IDF fallback โ zero API keys required.
Browser, Shell, GitHub, Gmail, Calendar, Drive, Notion, Trello, Twitter, Spotify, Apify, Image Gen, Knowledge Base, Scheduler, and more. Plus any external MCP server.
A 5-stage pipeline that turns natural language into validated, delivered outcomes.
Your natural-language goal is converted into a structured definition with typed, measurable success criteria โ file existence, content patterns, delivery receipts โ before a single action is taken.
Complex goals are broken into ordered sub-tasks with declared outputs. The right specialist agent is auto-selected by keyword or explicitly @mentioned.
Agents reason, pick tools, execute, and assess in a tight loop. Up to 5 sub-agents run in parallel. When a script fails, a semantic repair engine fixes it and retries.
After every iteration the system checks real outcomes. The agent cannot self-report "done" โ it's done when filesystem, content, and messaging evidence all pass.
Results delivered to Telegram, Discord, Slack, or your workspace. State is checkpointed โ a restart picks up exactly where it stopped, all artifacts intact.
Goals close only when real-world evidence passes โ no self-reporting allowed.
Fails a step? Replans, tries alternatives, semantically repairs scripts, retries.
State saved after each iteration. Restarts pick up exactly where they stopped.
Send goals, approve scripts, receive files โ all from Telegram, no UI needed.
Anthropic, Gemini, OpenAI, Ollama, LM Studio, Groq, DeepSeek, Perplexity.
Type @handle or just a keyword โ orchestrator routes to the right specialist.
Every tool call, file write, script run, and approval logged. Full run history.
Goals triggered by webhooks, Google Calendar events, or Gmail arrivals.
DALL-E 3 or Stability AI from inside any agent workflow.
Upload PDFs, URLs โ chunked, embedded, searched for grounded answers.
URL safety checks, path guards, AES-256-GCM secrets, optional JWT auth.
Points to local Ollama by default. Start free, add keys progressively.
One env var to activate any integration. Auto-discovered at startup.
Python + yfinance runs in Docker sandbox. Styled multi-sheet Excel with embedded chart. File criterion confirms output.
Scheduler fires โ GitHub data fetched โ report written โ Slack delivery. Set once, runs forever.
First run: Python from scratch. After success, system extracts a reusable template. Second run: template fills in โ instant execution.
V2 is the multi-tenant, governed edition. Deploy once for your company โ every member gets isolated agents, admins get full audit and control, and an opt-in AI proxy can act on your behalf, always behind your approval. Invite-only.
No scripts. No staging. AI Partner joins a live call, listens in real time, and responds.
Double-booked by default? AI Partner attends meetings you can't make, follows up on your behalf, monitors your market while you sleep. Your time is your revenue โ stop wasting it on tasks that don't need you.
Spend 6-8 hours daily in meetings most don't need your judgment โ just your presence. Send the proxy. It attends, flags blockers against your OKRs, captures decisions, and reports back in minutes.
AI Partner drafts replies to your emails, Slack DMs, and Telegram messages in your style. Nothing sends on its own โ every action on your behalf is gated by AuthorityPolicy and a one-tap approval, and AI involvement is always disclosed.
Joins your Meet/Teams/Zoom calls with AI presence disclosed, listens, takes notes, and contributes only what you've authorized โ then generates summaries and extracts action items automatically.
V2 doesn't wait to be asked. A background loop evaluates what to do every 15-30 minutes using your goals, memory, and real-time context.
From 17 servers in V1 to 36 in V2. Plus 4-tier computer use (DOM โ Vision โ Container Desktop โ Host Control) with live CAPTCHA handoff.
Screens calls for you, then records, transcribes, and summarizes. Places approved outbound calls from your Twilio number โ every call authorized by you, with AI disclosure on connect.
Episodic + Biographic + Counterparty + Vector + RAG. Understands who's who, what you've discussed, and maintains relationship context across all channels.
From 8 providers in V1. Now including NVIDIA NIM, Cerebras (~100k tokens/s), OpenRouter (100+ models), MiniMax, and LiteLLM Proxy.
Declarative rules define what the agent does automatically vs. what needs your approval. Telegram inline buttons for every gray-zone decision.
V2 is multi-tenant and governed from the core โ the reason it's safe to deploy across a whole company.
One install hosts your whole team. Every user's chats, files, memory, and screenshots are fully isolated โ no cross-user leak, enforced at the data, event, and identity layers.
A control room with a live activity feed, users & invites, an org audit log, per-user usage & cost, skill governance, knowledge sources, and agent access.
Every tool call, file write, and approval is logged. Declarative rules decide what runs automatically vs. what needs a human โ per action, per relationship.
No default logins. First registrant becomes admin; everyone else joins by single-use invite code. JWT auth, encrypted credential vault, AES-256 secrets.
Expose your agents to other systems over authenticated agent-to-agent calls. Service-account consumers, metered usage, CSV billing, hostile-agent detection.
Sync Notion, Slack, and Drive into a shared knowledge base on a schedule. Every member's agent answers grounded in your company's own documents.
The same platform, pointed at one job. Email in โ reconciled books โ P&L emailed back. This is what "AI that works for your clients" looks like in production.
Each client gets their own isolated tenant and a dedicated address. They email bank and credit-card CSVs โ attachments are pulled in automatically.
The bookkeeper profile categorizes transactions, reconciles against prior balances, and flags anomalies โ in a sandbox, building a live-formula workbook.
Before anything reaches the client, you get a one-tap human approval. Review the numbers, then release โ nothing goes out on its own.
A finished Excel workbook with live formulas and a zero-difference reconciliation lands back in the client's inbox. Month-end close, on autopilot.
Secure your spot now. Early access members get first invites, exclusive onboarding, and direct founder contact.
๐ฅ Spots are limited โ once beta fills, waitlist closes.
Single Docker Compose. One API key minimum โ or use local Ollama for free.
Book a walkthrough for your team, or get a service pilot scoped to one of your clients. Drop your email and we'll reach out.
๐ No spam โ we'll only reach out about your demo or pilot.
Self-hosted. Local-first. The relationship is the product โ and it's yours.