AI Engineer

Stop Rationing Tokens: Let the Harness Pick the Model — Kimchi by Cast AI

Published 2026-10-02 · Duration 18:08

Summary

Cast AI introduced Kimchi, an open-source coding harness designed to address the unsustainable cost of LLM tokens. Instead of limiting developers by cost per token, Kimchi shifts the focus to 'cost per task,' automatically selecting the optimal proprietary or open model for each step based on the required outcome. The platform provides a full software development lifecycle solution, including Ferment for multi-hour autonomous coding runs, Teleport for remote sandboxes, and Studio for team-based agentic collaboration.

Download summary

Key takeaways

  1. Cost Metric Shift: Task vs. Token 5:05

    The core principle is that comparing LLM models based solely on cost per token is misleading. Kimchi measures the true cost per task, allowing for automated model selection to maximize efficiency and minimize cloud expenditure.

  2. Autonomous Model Selection 7:10

    The Kimchi harness acts as an automated engine, selecting the best model (proprietary or open) for a given task at the right time, optimizing cost based on the desired outcome.

  3. Autonomous Development Workflow (Ferment) 13:35

    Ferment enables milestone-based, self-scoring autonomous coding runs that can deploy to staging. It constantly checks code quality, ensuring the output meets a minimum score (e.g., B or A) before proceeding.

  4. Remote and Continuous Workflows (Teleport)

    Kimchi Teleport spins up a secure sandbox in a remote environment (e.g., Google Cloud or on-premise Kubernetes cluster), allowing agent sessions to continue running even if the user closes their laptop or is in transit.

Technical details

  • Cost Optimization 449s

    The platform achieves significant cost savings (2.5x over three months of internal use) by optimizing model selection. The goal is to provide an 'essentially unlimited' and inexpensive coding agent.

  • Ferment Process 815s

    Ferment manages long-running tasks by breaking them into milestones. It incorporates a higher-order model that checks the quality of the code output, iteratively prompting the agent to fix issues until a target score is met.

  • Kimchi Teleport

    Teleport synchronizes the user's entire environment with a container running in a hyperscaler (like Google Cloud) or on-premise Kubernetes cluster, ensuring continuous agent operation regardless of local connectivity.

  • Kimchi Studio

    Studio is designed for team collaboration, visualizing and managing agentic tasks on a Kanban-style board. It allows multiple team members (including non-technical personnel) to review, take over, and collaborate on the agent's work-in-progress.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.