GitHub’s Post

View organization page for GitHub

6,680,091 followers

🆕 Anthropic's Claude Sonnet 5.5 is now generally available in GitHub Copilot. Designed for well-scoped everyday work like building features and fixing bugs, Claude Sonnet 5.5 stood out for its efficiency in our early testing: It matched Sonnet 5 on coding tasks while using significantly fewer steps, tokens, and tool calls. Plus, it finished tasks noticeably faster. ⚡️ Try it out in the GitHub Copilot app, CLI, and Visual Studio Code. https://lnkd.in/eHqZ35gW

Claude finally learned the ancient engineering principle: maybe don’t open twelve tabs.

The model cadence is wild. Positioning Sonnet 5.5 at well-scoped everyday work is smart, features and bug fixes are where devs spend most of their week, so that is where the compounding shows up fastest.

Like
Reply

Fewer steps on well-scoped work, nice. What does it do when the issue isn't well scoped: ask, or finish the wrong thing faster? 🤔

Like
Reply

Fewer steps and tool calls are promising if the resulting patches remain easy to review. It would be useful to compare failure recovery too: how often does the agent notice a failing test, correct its own change, and leave a concise explanation of what happened?

Like
Reply

Exactly. For me, fewer steps would make the biggest difference too. The less there is to review and double-check, the easier it is to trust the output and keep moving.

Like
Reply

Fewer tokens per task cuts costs in high-volume build pipelines. For automation running thousands of times, that margin compounds. Speed in iteration cycles is where it counts.

Like
Reply

On the underspecified-task question: I'd include deliberately ambiguous issues in the eval, with “ask a clarifying question” as a valid outcome. Score whether the acceptance criteria were met, not just steps or tokens. Otherwise a model can look more efficient by skipping the question that would prevent the wrong change.

Like
Reply

Fewer steps is the number I'd look at first. I'm not a developer, so most of my time goes to reading back what Claude did and checking it did what I asked. The shorter that trail, the sooner I can trust it and move on!

Like
Reply
See more comments

To view or add a comment, sign in

Explore content categories