# Stop Rationing Tokens: Let the Harness Pick the Model — Kimchi by Cast AI

## Executive summary

Cast AI introduced Kimchi, an open-source coding harness designed to address the unsustainable cost of LLM tokens. Instead of limiting developers by cost per token, Kimchi shifts the focus to 'cost per task,' automatically selecting the optimal proprietary or open model for each step based on the required outcome. The platform provides a full software development lifecycle solution, including Ferment for multi-hour autonomous coding runs, Teleport for remote sandboxes, and Studio for team-based agentic collaboration.

## Key takeaways

- Cost Metric Shift: Task vs. Token: The core principle is that comparing LLM models based solely on cost per token is misleading. Kimchi measures the true cost per task, allowing for automated model selection to maximize efficiency and minimize cloud expenditure.
- Autonomous Model Selection: The Kimchi harness acts as an automated engine, selecting the best model (proprietary or open) for a given task at the right time, optimizing cost based on the desired outcome.
- Autonomous Development Workflow (Ferment): Ferment enables milestone-based, self-scoring autonomous coding runs that can deploy to staging. It constantly checks code quality, ensuring the output meets a minimum score (e.g., B or A) before proceeding.
- Remote and Continuous Workflows (Teleport): Kimchi Teleport spins up a secure sandbox in a remote environment (e.g., Google Cloud or on-premise Kubernetes cluster), allowing agent sessions to continue running even if the user closes their laptop or is in transit.

## Technical details

- Cost Optimization: The platform achieves significant cost savings (2.5x over three months of internal use) by optimizing model selection. The goal is to provide an 'essentially unlimited' and inexpensive coding agent.
- Ferment Process: Ferment manages long-running tasks by breaking them into milestones. It incorporates a higher-order model that checks the quality of the code output, iteratively prompting the agent to fix issues until a target score is met.
- Kimchi Teleport: Teleport synchronizes the user's entire environment with a container running in a hyperscaler (like Google Cloud) or on-premise Kubernetes cluster, ensuring continuous agent operation regardless of local connectivity.
- Kimchi Studio: Studio is designed for team collaboration, visualizing and managing agentic tasks on a Kanban-style board. It allows multiple team members (including non-technical personnel) to review, take over, and collaborate on the agent's work-in-progress.

## Practical implications

- Shifts the focus of AI development from managing token budgets to optimizing task outcomes.
- Enables continuous, background agentic workflows that are not dependent on the user's physical location or device status.
- Provides a structured, collaborative environment for teams to work with AI agents, making agentic engineering accessible to non-technical stakeholders.

## Topics

AI Agents, LLM Cost Optimization, Agentic Engineering, Software Development Lifecycle (SDLC), CI/CD, Cloud Computing, Kimchi, Kimchi (GitHub), Cast AI

Source: https://www.youtube.com/watch?v=48YUYDjwfYY
