# Is Opus 5.5 Better At Writing?

## Executive summary

This evaluation compares the writing capabilities of Opus 5 and Opus 5.5 across six diverse tasks (FAQ generation, summarizing technical posts, drafting replies, and writing abstracts). While Anthropic claims Opus 5.5 writes more naturally and follows instructions better, the evaluation concludes that both models still produce 'slop'—content that requires significant human editing. Opus 5.5 generally performed better, particularly in structure and flow, but the overall finding is that AI-generated content, especially in complex technical writing, is far from ready for deployment without heavy refinement.

## Key takeaways

- Opus 5.5 shows structural improvements over Opus 5.: In multiple blind comparisons, Opus 5.5 was observed to be better, particularly in organizing information and maintaining a more natural flow. However, the difference was not always definitive, and both models struggled with complex, nuanced writing tasks.
- Models frequently over-correct or hallucinate details.: When summarizing technical posts, models tended to over-steer based on the prompt, adding unnecessary context or fabricating details (e.g., inventing the 'spam' example for a classifier). This highlights a risk of factual inaccuracy in AI output.
- The challenge of 'Slop' persists.: Despite improvements, the consensus is that 'slop' (low-quality, generic, or overly structured content) is not solved. The models often write the same thing in the same order, making the output difficult to trust without rigorous human review.
- The best performance was seen in structured, constrained tasks.: The models performed best when given highly specific context or when the task was limited (e.g., writing a short FAQ). However, even in these cases, the output required substantial editing.

## Technical details

- Model Comparison: The evaluation compared Opus 5 and Opus 5.5 across six tasks: FAQ generation, summarizing a blog post (on Jev), drafting a thoughtful technical reply, writing a LinkedIn influencer post, drafting a short reply for engagement, and creating a conference talk title/abstract.
- AI Evaluation (Evals): The concept of 'evals' was introduced as a process to measure an AI application's performance, noting that it is usually not a single test or score. The speaker recommends a free guide for building evals.
- LLM Judge and Classifiers: The discussion touched upon the use of LLM judges and classifiers for measuring quality. One suggested improvement was to prompt the model to 'Measure your LLM judge like a classifier' rather than simply asking it to measure quality.
- Technical Identifiers: Models evaluated include Opus 5, Opus 5.5, Gemini 3.5 flash, and the judgment model Jev. The tasks involved analyzing concepts like Retrieval-Augmented Generation (RAG) and LLM pipelines.

## Practical implications

- Do not rely on AI output for mission-critical or highly nuanced content without extensive human review and editing.
- When building AI products, focus on developing robust evaluation pipelines (Evals) that test specific failure modes rather than relying on general quality scores.
- Prompt engineering must be highly specific, guiding the model not just on content, but on tone, structure, and the required level of detail to prevent 'slop'.

## Topics

Large Language Models (LLMs), Prompt Engineering, AI Evaluation (Evals), Natural Language Processing (NLP), Generative AI, Free Evals Guide

Source: https://www.youtube.com/watch?v=8uafNYFJ0rU
