Skip to main content

OrcA 0.4: An MCP Agent Harness in VS Code, for Any MCP Server

· 14 min read
Adrian Escutia
La Rebelion Founder

OrcA 0.4 turns VS Code into an agent harness for MCP and API work. Connect any Streamable HTTP MCP server, not only HAPI ones. Let the agent load tools as the conversation moves instead of drowning in 40 schemas per turn. Read and edit exactly what the model is told, and see the goals, facts and plans behind a multi-step request. Still free, still no telemetry (by default).

TL;DR​

  • What: OrcA 0.4, a free VS Code extension to harness MCP and API work.
  • The shift: 0.3 showed you what your agent did. 0.4 shapes what it receives: instructions, tools and context. That is the job of a harness.
  • New: External MCP servers (any Streamable HTTP server), progressive tool discovery, the Harness tab, the Context tab, and Generate Contract for hand-written servers.
  • Why it matters: fewer tokens per turn, fewer wrong tools, no more "I don't have access to that tool" when the tool is right there.
  • Install: search OrcA MCP in VS Code Extensions, or code --install-extension la-rebelion-labs.orca-mcp.
OrcA 0.4 walkthrough: external MCP servers, progressive tool discovery, and the Harness and Context tabs.

The problem: the agent can't see the right tool​

OrcA 0.3 gave you the traffic graph, and it showed a pattern we kept seeing in real sessions:

  • The model receives every tool schema on every turn. With a few servers, that is thousands of tokens before the user types a word.
  • Accuracy drops as the list grows. Similar tools (book vs. reschedule) get mixed up.
  • Hide tools to save tokens, and the agent finds the first step, then announces "I don't have access to a tool to create the appointment", because it never learned the booking tool existed.
  • The instructions that drive all of this live in a setting nobody opens.

These are not model problems. They are harness problems: what the model is given, when, and under which rules.

What is an agent harness?​

A harness is the software layer that runs an agent session. It prepares the instructions, context and tool definitions for the model, applies approvals, routes tool calls to the right place, and records what happened. VS Code ships general-purpose harnesses for coding.

OrcA is a specialised harness for MCP and API work. It does not edit your files or type into your terminal. It connects models to MCP servers, decides which tools they see, enforces confirmations, explains business context, and shows every exchange. It is where you make sure your tools are agent-ready before Copilot, Claude, or an AI teammate like Clooney from Clawne Me uses them.

What's new in OrcA 0.4​

🔍 Progressive tool discovery: tools arrive when the conversation needs them​

When a server offers a search tool (HAPI's tool_search, or any tool_search_tool*), OrcA no longer sends every definition. Instead:

  • The model gets a small search_tools tool and, inside its description, a catalog of tool names that are not loaded yet. It knows bookAppointment exists before it needs it, even when the API speaks another language than the user.
  • Before the model answers each message, OrcA searches with your message and loads the matching tools. "Now book it" brings the booking tool.
  • The model can also search by intent, or load exact tool names from the catalog.
  • Discovered tools are appended, never reordered, so provider prompt caches keep working. New Chat resets them.
  • Chat cards say who loaded what: search_tools, Tools loaded for your request, or Tools loaded from capability result.

MCP Servers marks progressive servers and their deferred tools; hover a server to read its Discovery line, for example Progressive tool discovery: tool_search (8 of 11 tools deferred). Choose the behavior with orca.chat.toolDiscovery: auto (default), always (OrcA searches any server locally, even without a search tool), or off (0.3 behavior).

🧰 The Harness tab: see and edit what the model is told​

The OrcA view gets a Harness tab next to Chat, Context and Activity:

  • Instructions: edit the chat's system prompt and save it for you or for the workspace; Reset brings back the default.
  • Automatic guidance: read, word for word, what OrcA adds (for example "search before saying a tool is unavailable").
  • Settings: discovery mode, search limit, tool rounds, and the confirmation policy.
  • Token preview per server: tools sent upfront vs. the full catalog. This is the context you save, as a number.

No hidden prompts. If the agent behaves strangely, the first place to look is now one click away.

🌐 External MCP servers: bring any Streamable HTTP server​

OrcA is no longer HAPI-only:

  • Add External MCP Server…: a name, a Streamable HTTP URL, optional headers.
  • Import MCP Servers… from .vscode/mcp.json and your user mcp.json (stdio and SSE entries are listed as not supported).
  • External servers appear in a new External group with Connect, chat, Traffic, Expose via Proxy, Add to VS Code, Edit and Remove. Headers stay in VS Code's secret storage.
  • A server that answers 404/405 is reported as not speaking Streamable HTTP, instead of failing silently.

📄 Generate Contract: a contract for hand-written servers​

Built your server by hand? Generate Contract infers a deterministic OpenAPI 3.1 document from its tools, resources, resource templates and prompts. It is linked to the server, so Open Operation in Contract, Traffic details and chat tool cards work as they do for HAPI servers, and Dry Run works on it.

When the server's tools change, it shows contract out of date; Regenerate Contract opens a diff and writes only when you choose Save Changes.

🧭 The Context tab: goals, facts and plans​

On HAPI servers served with --capability-graph --capability-planning, the Context tab shows the business map behind a request:

  • With each message, OrcA asks capability_context (read-only) and lists the candidate goals, the facts each step requires, and the gates, such as a confirmation. The tools it names are loaded for the model.
  • Plan This Goal asks capability_plan with the facts you have verified, and shows the steps, the missing facts and the unsatisfied gates.
  • From Chat keeps the results the conversation produced, with Show in Traffic.

Nothing is executed from this tab. It answers "what has to be true before the agent can book this?" in plain view.

✨ Smaller things that add up​

  • Activity now records what you do with MCP servers: adding, importing, editing and removing servers; connections (with the catalog size or the failure reason); header changes (names only); contract generation; proxy exposure; and New Chat.
  • Clearer HAPI API errors: when the HAPI API rate-limits or fails, OrcA names the API and the mode, retries after a rate limit, offers Switch to Local Mode, and keeps your local and external servers listed.
  • Filter Contracts… filters the Contracts view as you type.
  • Dry Run opens as a preview tab in the current editor group.

OrcA 0.4 compared with 0.3​

QuestionOrcA 0.3OrcA 0.4
Which servers can I use?HAPI serversAny Streamable HTTP MCP server, added or imported
Which tools does the model receive?All of them, every turnOnly what the conversation needs, plus a catalog of names
Can the agent find a tool for the next step?Only if it was already sentLoaded per message, by search, or from capability context
What is the model told?Hidden in a settingHarness tab: instructions, automatic guidance, settings
What must be true before a multi-step action?Not visibleContext tab: goals, facts, gates, plans
Is there a contract for a hand-written server?NoGenerate Contract (OpenAPI 3.1), with drift detection

How to give an agent progressive tools in VS Code (3 minutes)​

  1. Install OrcA from the VS Code Marketplace.
  2. Add a server: in MCP Servers, choose Add External MCP Server… (or Run with HAPI on an OpenAPI file, served with a search tool).
  3. Connect: right-click the server → Connect, then hover it to read the Discovery line.
  4. Tune the harness: open OrcA → Harness, check the instructions and compare the tokens sent upfront with the full catalog.
  5. Chat: ask a multi-step question, such as "find a pediatrician next Tuesday, then book the first slot". Watch the cards load the right tools at each step, the Context tab explain the goal, and Traffic record every call.

Make your own APIs harness-friendly​

The harness can only be as good as the contracts behind it. If you own the API, two guides show how to make it agent-ready with HAPI, no code required:

Private by design, as before​

  • No telemetry by default and no vendor AI SDKs.
  • The chat only contacts the provider you selected, or VS Code's own models.
  • API keys and server headers live in VS Code's secret storage.
  • Traffic stays in memory with secrets redacted; nothing is written to disk until you save a session.

Good to know​

  • OrcA connects to Streamable HTTP servers only; stdio and SSE are not supported.
  • OAuth-protected servers are still not supported; static headers work.
  • HAPI servers behave as in 0.3; orca.chat.toolDiscovery: off restores the 0.3 tool list exactly.

Frequently asked questions​

What is an agent harness? The layer that runs an agent session: it prepares instructions, context and tools for the model, applies approvals, routes tool calls and records what happened. OrcA 0.4 is a harness specialised for MCP and API work; it does not edit files or run a terminal.

What is progressive tool discovery in MCP? Sending a small search_tools tool and a catalog of tool names instead of every definition, then loading full definitions as the conversation needs them. OrcA searches with each message, appends discovered tools to keep prompt caches valid, and resets them on New Chat.

How do I reduce MCP context bloat? Serve your API with a search tool (for HAPI, --deferred-tool-search or --capability-graph) and keep orca.chat.toolDiscovery on auto, or use always to let OrcA search any server. The Harness tab shows the tokens you save.

Can OrcA connect to MCP servers not built with HAPI? Yes: Add External MCP Server… or Import MCP Servers… from mcp.json, for any Streamable HTTP server. stdio and SSE are not supported.

Can I get an OpenAPI contract for a hand-written MCP server? Yes: Generate Contract infers an OpenAPI 3.1 document from the server's catalog, links it to the server, and flags it when the tools change.

Does OrcA 0.4 support OAuth-protected MCP servers? Not yet. Static headers such as a bearer token or an API key work today and are stored in VS Code's secret storage.

Get OrcA 0.4​

0.3 let you see the conversation between your agent and your tools. 0.4 lets you shape it.

Be HAPI, and enjoy building!