• Agent Pulse
  • Posts
  • OpenAI Just Called the Coding-Agent Race Rigged

OpenAI Just Called the Coding-Agent Race Rigged

Grok Enters the Office as Coding Benchmarks Crack

In partnership with

You Already Have a Take on What AI Does Next

OpenAI or Anthropic? Which model leads the next benchmark? Which company ships the next major breakthrough?

If you follow AI closely, you already have opinions on where the industry is headed. Kalshi lets you trade on real-world AI and technology events, with markets that move as models launch, benchmarks drop, and announcements happen.

The people who follow this space most closely often see the story developing before everyone else. Put that knowledge to work and trade what you think happens next.

Bonus credit varies from $15 to $500. Terms apply.

Welcome back! OP here again, helping you with another addition of Agent Pulse - your go-to spot for agentic news, insights and more.

In today’s:

  • 👉 TOP Agentic News

  • ✨ Featured Agents

  • 🗺️ Agents Landscape Map

OpenAI published an audit estimating that about 30% of SWE-Bench Pro tasks are broken. SWE-Bench Pro is used to measure longer-horizon coding-agent capabilities, and OpenAI noted that frontier model performance on the public split rose from 23.3% to 80.3% in eight months. The concern is that flawed evals can distort product claims, deployment decisions, and research priorities. OpenAI’s review used model traces, metadata, investigator-agent passes, and experienced software engineer review. Strategically, the coding-agent market needs cleaner, harder, contamination-resistant evals before buyers treat benchmark jumps as real product truth.

xAI launched Grok 4.5, positioning it for coding, agentic workflows, and office productivity. In Grok Build, xAI says it can create complex Excel models using web research, multi-sheet formulas, and structured analysis. It also supports PowerPoint and Word workflows using native shapes, diagrams, slide design, and prose generation. Pricing is listed at $2 per million input tokens and $6 per million output tokens, with xAI claiming roughly 2x token efficiency versus comparable leading models. Strategically, model competition is moving from raw benchmark claims toward cost-per-completed workflow inside real productivity tools.

See the whole platform. No guided tour.

Skip the sales call. Walk through Gladly's interface yourself — the AI suggestions, the unified customer view, the full conversation thread. 15 minutes, no installation, no commitment.

Accela acquired Civira AI, an AI-agent platform built for civic technology. Civira agents automate configuration from documents and forms, generate documentation, author and test scripts, build role-based apps, and answer configuration questions. Accela plans to embed the technology across deployment, implementation, and maintenance of government software. The goal is to reduce the time, cost, and technical burden of modernizing permitting, licensing, and code-enforcement systems. Strategically, govtech agent adoption may start in back-office implementation before it reaches citizen-facing services.

Akeneo has introduced Agentic Ziggy, a new agentic orchestration layer embedded directly within the Akeneo Product Cloud . This release transforms the platform from a traditional system of record into a fully agentic system of action, where fleets of specialized AI agents coordinate to manage product data. These agents handle complex tasks such as data modeling, schema mapping, and continuous quality checks, significantly reducing the manual workload for catalog teams.

The platform also features prompt-based AI image editing and intelligent error management for syndication, allowing teams to resolve issues at scale through self-service workflows. By automating these operational complexities, Akeneo is positioning brands and retailers to thrive in the emerging era of agentic commerce and discovery.

Abrigo announced Abrigo Agentic Platform Experience, or APX, for financial institutions. The platform is designed to automate, orchestrate, and scale work across growth and risk-management functions. Its lending launch is planned for Q3 2026 and covers the loan lifecycle from pipeline management and underwriting through closing, servicing, and portfolio administration. Abrigo estimates agentic AI can reduce manual labor by more than 40% while improving borrower experience. Strategically, regulated financial agents are being packaged around explainability, governance, and human oversight from day one.

200+ Proven Ways to Make Money With AI in 2026

The next wave of millionaires will be people who figured out how to make AI work for them.

The window to get ahead is still open. But not for long.

Here are 200+ proven ways to make money with AI in 2026.

Sign up for Superhuman AI, the free daily newsletter read by 1M+ professionals, and get instant access to all 200+ ways to profit from AI this year.

Featured AI Agents
  • Loopa - AI agent platform for workflow automation and content creation

  • Imagvio AI - AI image editor and video creation platform with character consistency and Nano Banana model.

  • eSkilled - Powerful Online Course Maker, Build High Quality Courses 100x Faster!

  • Handinger - Turn any website into clean Markdown

  • Freeday - Enterprise platform for deploying AI agents

  • Agentman - Bring joy back to practicing medicine

  • IBBS - A security badge proving your AI agent resists prompt injection attacks

  • Soulmate IO - Private AI companions with voice chat and persistent memory

  • Nano Banana - Prompt-based photo editing with character consistency

🗺️ The Map of AI Agents (Live & Growing)

We’re charting the entire AI agent ecosystem — thousands of options across categories.
Your next agent is already on the map.

THANK YOU 

Follow us on

I appreciate your time.

OP & Team

How'd we do?

Login or Subscribe to participate in polls.

Reach 23,000+ Readers:

Newsletter is read by VCs, founders, engineers, managers and tech professionals.

Reply

or to participate.