Hamel Husain

Is Opus 5.5 Better At Writing?

Published 2026-10-08 · Duration 19:56

Summary

This evaluation compares the writing capabilities of Opus 5 and Opus 5.5 across six diverse tasks (FAQ generation, summarizing technical posts, drafting replies, and writing abstracts). While Anthropic claims Opus 5.5 writes more naturally and follows instructions better, the evaluation concludes that both models still produce 'slop'—content that requires significant human editing. Opus 5.5 generally performed better, particularly in structure and flow, but the overall finding is that AI-generated content, especially in complex technical writing, is far from ready for deployment without heavy refinement.

Download summary

Key takeaways

  1. Opus 5.5 shows structural improvements over Opus 5.

    In multiple blind comparisons, Opus 5.5 was observed to be better, particularly in organizing information and maintaining a more natural flow. However, the difference was not always definitive, and both models struggled with complex, nuanced writing tasks.

  2. Models frequently over-correct or hallucinate details. 5:37

    When summarizing technical posts, models tended to over-steer based on the prompt, adding unnecessary context or fabricating details (e.g., inventing the 'spam' example for a classifier). This highlights a risk of factual inaccuracy in AI output.

  3. The challenge of 'Slop' persists.

    Despite improvements, the consensus is that 'slop' (low-quality, generic, or overly structured content) is not solved. The models often write the same thing in the same order, making the output difficult to trust without rigorous human review.

  4. The best performance was seen in structured, constrained tasks. 0:55

    The models performed best when given highly specific context or when the task was limited (e.g., writing a short FAQ). However, even in these cases, the output required substantial editing.

Technical details

  • Model Comparison 0s

    The evaluation compared Opus 5 and Opus 5.5 across six tasks: FAQ generation, summarizing a blog post (on Jev), drafting a thoughtful technical reply, writing a LinkedIn influencer post, drafting a short reply for engagement, and creating a conference talk title/abstract.

  • AI Evaluation (Evals) 55s

    The concept of 'evals' was introduced as a process to measure an AI application's performance, noting that it is usually not a single test or score. The speaker recommends a free guide for building evals.

  • LLM Judge and Classifiers

    The discussion touched upon the use of LLM judges and classifiers for measuring quality. One suggested improvement was to prompt the model to 'Measure your LLM judge like a classifier' rather than simply asking it to measure quality.

  • Technical Identifiers 223s

    Models evaluated include Opus 5, Opus 5.5, Gemini 3.5 flash, and the judgment model Jev. The tasks involved analyzing concepts like Retrieval-Augmented Generation (RAG) and LLM pipelines.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.