Subscribe Sign in

AWS bolts together open source agent harness, says it sips fewer tokens than rivals

1 min read Rewritten in plain language

Artificial intelligence

Show what we removed Rules applied: A6 C2 D2×2 D3×12 E3×4 F2×2 all 30 rules
  • Strands claims near-parity with Claude Code and Codex on benchmarks, though it only moved coding agents and marked its own homework.
  • On top of being relatively plug-and-play in design, AWS claims the Strands harness achieved "nearly equal benchmark scores" versus Claude Code, Codex, and “other popular harnesses” when tested using the Harbor framework, distributed across multiple nodes of AWS’ own EC2 virtual servers for benchmarking tests.
  • Per the announcement post, it defaults to truncating tool results over 1,500 tokens, automatically compacting its context window when it surpasses 85 percent, and automatically trying to recover context in overflow cases.

3 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

Headline check

There is nothing in this headline a machine can check against the report: no figure, no name and no quotation.

Nothing was measured here, so nothing is claimed. How this is checked

Strands claims near-parity with Claude Code and Codex on benchmarks, though it only moved coding agents and marked its own homework.

AWS has entered the open source agentic AI game, claiming its new Strands harness matches rivals on benchmarks while using around a quarter fewer tokens.

The Strands harness, as its name suggests, is built on AWS’ Strands Harness SDK, but is packaged up and ready to roll out of the box, either locally or deployed to work with whatever AI provider a customer prefers.

On top of being relatively plug-and-play in design, AWS claims the Strands harness achieved "nearly equal benchmark scores" versus Claude Code, Codex, and “other popular harnesses” when tested using the Harbor framework, distributed across multiple nodes of AWS’ own EC2 virtual servers for benchmarking tests.

Strands consumed 28 percent fewer tokens across six benchmark tests when compared to “Claude or GPT models,” claims AWS, and in some cases had better accuracy than other harnesses too.

AWS credits this performance to the Strands harness’ default prompt caching and context management settings. Per the announcement post, it defaults to truncating tool results over 1,500 tokens, automatically compacting its context window when it surpasses 85 percent, and automatically trying to recover context in overflow cases.

Shortened to 1 minute of reading, this version reads 5.7 on the Niral Score.

You are reading our version, not theirs. This is The Register's report shortened to its most important sentences, in plainer words, with verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingHappiness
The Registeras they published this story 8.6 9 69 47.9
Mundane Readneutralized from The Register 8 9 69 47.9

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works