Skip to content

Cake Agents Success Story: Own Your Agent Harness

Published: 10/2026
7 minute read
illustration of a workflow with agents

Cake Agents enabled an engineering organization to cut token costs 50%, protect PII and proprietary data, and double agentic engineering productivity, all while keeping the harness running in their own cloud environment.

Prior to installing Cake Agents, the customer was using Cursor and Claude Code heavily. The agents worked, but local execution created a new bottleneck: laptop resources, git worktrees, and human attention. While the existing vendors did offer scalable cloud-based alternatives, these did not meet the customer’s needs for security and data privacy (including a needed Zero Data Retention policy).

The customer moved to Cake Agents to unlock the scale and security of the cloud, while maintaining full control over the agentic coding environment and all of their data. By moving most of his agentic development to Cake Agents, it was possible to double session concurrency, expand agent-assisted review, and gain control over model selection and inference economics.

Key takeaways

  1. 1

    Cut inference costs >50% by deploying Cake-managed LiteLLM integrated with Cake Agents, to unlock "bring your own model" and Fireworks for inference.

  2. 2

    Maintained control over data and infrastructure by deploying Cake Agents inside the customer cloud with Zero Data Retention by default.

  3. 3

    Doubled agentic engineering concurrency for IC, by running 3–5 implementation agents and up to 3 additional agent-guided review sessions simultaneously.

  4. 4

    After a small scale trial, the customer is rolling out Cake Agents to an 80-person engineering team and is beginning to automate workflows such as DataDog alert → Cake Agent → pull request.

  5. 5

    During testing, the customer completed a 14,000-line architectural rework in under 4 hours with >95% of the work done in Cake Agents sessions.


Moving coding agents off the laptop

The company’s engineering director adopted Cake Agents as an early adopter. He works across a broad technical surface area: full-stack product development, agent harnesses, and underlying cloud infrastructure.

Like many engineers adopting coding agents, he started primarily with local tools such as Cursor and Claude Code. But as agents took on more substantial engineering work, the laptop itself became a bottleneck.

Running several local agents simultaneously meant maintaining multiple worktrees while each session competed for the same compute resources. In practice, managing more than three concurrent local sessions became difficult.

Cake Agents moved those workloads into parallel cloud environments instead.

“My concurrency roughly doubled or better, and the ceiling stopped being my laptop and started being how fast I can review and steer.”

Today, he typically runs 3–5 Cake Agent sessions concurrently for implementation, sometimes adding another 3 sessions for agent-guided code review. That means as many as 8 active sessions at once.

Cursor and Claude Code remain part of his workflow, but increasingly at the edges: very small tasks such as running scripts or making one-line changes, and highly involved work requiring constant validation across the full stack.

The bulk of day-to-day engineering work has moved to Cake Agents.

A 14,000-line architectural change in under 4 hours

The advantages of background agents become particularly apparent on changes that are conceptually straightforward but mechanically large.

The team recently undertook a significant architectural rework of one of its core systems. They were moving from a heavily partitioned design to a consolidated architecture while keeping both implementations available behind a feature flag during rollout.

The resulting change touched approximately 14,000 lines of code, including new packages, restructured prompts and instructions, and extensive test coverage.

Cake Agents produced the bulk of the implementation in under four hours.

For an engineer working interactively with a local agent, a task of that scale can monopolize a workstation for hours or stretch across multiple days. Running it as a background Cake Agent changed the calculus of whether to take on the work at all.

“That’s the kind of change that’s mechanically large but conceptually coherent, exactly where a background agent shines.”

Instead of breaking the project into smaller pieces because of local resource constraints, the team could give the complete task to an agent, let it work independently, and focus human attention on steering and reviewing the result.

Bring your own models

Horizontal scalability was only one consideration in the company’s adoption of Cake Agents. The team also wanted to retain control over the infrastructure and economics behind its coding agents.

Cake-managed LiteLLM integrates with Cake Agents and enables the customer to bring their own model. Switching the core inference provider to Fireworks via LiteLLM cut token costs by approximately 50%, while improving code quality and speed of token generation.

Running Cake within the company’s own infrastructure provides another layer of control and data protection.

“Bring your own model and running in your own infra is huge. It means you can tune cost and model choice yourself as usage grows instead of having to go renegotiate with a vendor every time. There is also never any concern about controlling PII or proprietary data.”

For teams with security, data, or compliance requirements, keeping control of the underlying infrastructure, by running Cake Agents in their own cloud, removes an entire category of concerns from the agent-adoption process.

Using agents to review agents

The team also uses Cake Agents for user-guided agentic code review, alongside automated Claude review in CI.

Automated review catches many issues, but experienced engineers often have specific concerns based on context. Instead of pulling down a branch, navigating a large diff, and tracing the implementation manually, a reviewer can simply launch a Cake Agents session and investigate the change directly.

The human decides what matters and what to ask; the agent traces those concerns through the codebase and finds the relevant implementation.

This creates a middle ground between fully automated review and manually reading every line of a large pull request—without requiring reviewers to repeatedly pull and configure branches locally.

From pair programming to managing parallel agents

The largest gains from coding agents come from changing the workflow, not just speeding up individual tasks. Instead of working sequentially with one AI pair programmer, the team uses Cake Agents to delegate, monitor, review, and steer multiple tasks in parallel.

“The value really shows up when you change how you work, not just when you swap out the tool.”

That shifts the engineer’s role from primarily writing code to defining work, keeping multiple streams moving, reviewing results, and applying judgment where it matters most.

“The win is running several at once and getting good at reviewing and steering instead of typing.”

Work that was previously serialized across one or two local sessions can now happen concurrently, turning some multi-day changes into a single sitting.

What comes next

Having completed a pilot with a small number of users, the customer is now rolling out Cake Agents to an 80-person engineering team. Cake Agents supports secure multiplayer in the cloud, which has unlocked new collaborative agentic coding workflows that will be extended to the broader team.

Additionally, the customer has begun implementing automated workflows, such as initiating Cake Agents sessions to fix bugs based on DataDog alerts, in many cases going from alert to PR with no human effort.

The expanded deployment of Cake Agents is targeted at improving agentic engineering adoption and velocity while maintaining the security and scalability advantages of running Cake Agents in their own cloud. Additionally, keeping inference costs down is becoming increasingly critical as token consumption continues to grow, driven by higher parallelism, more users, and an increasingly automated software development lifecycle.