Computer-use agent for software without an API
A perceive-decide-act loop that drives legacy desktops and web apps through the screen, with approval gates.
Use it when
Systems that have no API or export: legacy ERPs, government portals, partner sites, thick-client tools. Use APIs and MCP wherever they exist; use the screen only where they do not.
The parts, top to bottom
across every level
Hover or tap any part to see what it does. The light shows the order a request moves through.
What happens, in order
- 1Task and policy defined
- 2Screenshot captured
- 3Vision model reads the screen
- 4Click / type / scroll action emitted
- 5Sandboxed execution & verification
- 6Approval gate before irreversible steps
- 7Trace with screenshots
What we typically build it with
Trade-offs
Top agents now exceed the 72% human baseline on OSWorld's desktop tasks, but long-horizon reliability is still the constraint, so we keep tasks short, verify state after each step and prefer an API the moment one appears.
Further reading
Problems this architecture solves
Generic problem statements with the flow and the outcomes the industry has documented.
Swivel-chair work across systems
Onboarding, reconciliation, procurement and approvals require people to read one system and type into another, with rules that live in someone's head.
How the system works
- Trigger (email, form, event)
- Orchestrator plans steps
- Worker agents read & decide
- Confirm risky actions
- Update systems & log
Outcome: CRM-native agents at a regulated vendor deflected 72% of support cases to self-service and saved 7.5 hours per case handled with AI assistance. Smarsh on Salesforce Agentforce
ArchitectureSoftware that has no API
A legacy ERP, a partner portal or a government site sits in the middle of the process, and the only way in is a screen and a keyboard.
How the system works
- Task & policy
- Screenshot
- Vision model decides
- Click / type in a sandbox
- Verify state
- Approve irreversible steps
Outcome: Top computer-use agents now exceed the 72% human baseline on OSWorld's real desktop tasks; kept to short, verified steps they remove the last manual hop in a process. OSWorld benchmark
ArchitectureOther agentic patterns
Single agent with tools
One model, a toolbox, and a loop that decides what to call next.
Use it when: A conversation or task needs judgement about which system to consult and in what order: customer queries, sales qualification, commerce in chat, internal copilots.
Flow
- 1User or event
- 2Agent reasons about the goal
- 3Calls a tool (search, CRM, calendar)
- 4Observes result
- 5Repeats until done
Supervisor multi-agent system
An orchestrator plans; specialised workers execute; results are merged.
Use it when: Multi-step processes that touch several systems or skills: research and due diligence, back-office case handling, candidate pipelines, anything where one agent's context would overflow.
Flow
- 1Task arrives
- 2Supervisor decomposes into steps
- 3Workers run in parallel with their own tools
- 4Results validated & merged
- 5Human checkpoint if needed
Let's build intelligent systems that drive growth
Tachyon is the engineering partner for teams that need AI in production, not in a deck. Start with a free 60-minute discovery call.
Tell us the problem, we will map it to the architecture
Every great partnership begins with a conversation. Whether you are exploring possibilities or ready to scale, tell us what you are actually trying to build.
- ubheshubham.37@gmail.com
- +91 84592 96471
- Clients worldwide · English
- Pune, India · Headquarters