SpaceRock

3 New Frontier AI Models: OpenAI, Meta and xAI make breakthroughs

By Veer Solanki · · 967 words

Topics: AI, Agentic AI, AI Agents, AI Arms Race, AI Benchmarks, AI Coding, AI Development, AI Image Generation

3 New Frontier AI Models: OpenAI, Meta and xAI make breakthroughs

Blinked this week?

You missed a generation. Just 72 hours ago, OpenAI announced its GPT-5.6 family in general release, SpaceXAI announced Grok 4.5 – its first model developed in parallel with Cursor – and Meta’s Superintelligence Labs jumped into the image generation market with Muse Image.

You’d have heard about each of these on their own.

But combined, they hint at something much larger: the AI frontier arms race isn’t about building the most intelligent AI – it’s about delivering it faster, cheaper, and at every scale.

OpenAI’s GPT-5.6 family rolled out from restricted preview to widespread release Thursday across ChatGPT, Codex and the API. It consists of three models by capability: Sol (its latest and greatest), Terra (its balanced, middle-of-the-road option) and Luna (its fastest, least expensive tier).

The numbers signify generations, and stable tiers like Sol, Terra and Luna signal capability tiers.

What these numbers signal isn’t just about intelligence but the efficiency and relative cost of each capability tier. Sol, OpenAI states, offers state-of-the-art performance on benchmarks across coding, cybersecurity and science, using fewer tokens than previous generations and its competitors.
For example, Sol outperformed Anthropic’s Claude Fable 5 by 13.1 points in the Agents’ Last Exam – a benchmark for long-horizon work across 55 professions – with OpenAI’s figures of 53.6 vs 40.5.

Its lower-capability counterparts will further drive down costs. With Terra and Luna, OpenAI claims “comparable or better results” to competitor models “at a fraction of the price.” The two new key feature options include: max mode, where the model can take more time to verify and refine its responses; and ultra mode, a four-agent mode for parallel, fast work on more demanding tasks at a higher token cost. With the addition of Programmatic Tool Calling – where the model can program and call tools with a few prompts without constant back and forth – GPT-5.6 is poised to power agents, not chatbots.

Here’s the pricing for GPT-5.6: $5 per million input tokens, $30 per million output tokens for Sol; $2.50/$15 for Terra; $1/$6 for Luna – with a 30-minute minimum retention window on its improved prompt caching system.

OpenAI highlights “most aggressive safeguards to date” for the cyber-capable GPT-5.6 and states that the latest version blocks nearly 10x more harmful activity than older versions and restricts access to its most sensitive capabilities via “Trusted Access,” which requires a verified identity. Access to models with the highest cyber capabilities will be restricted to hardware-backed passkeys starting in September, following 700,000 GPU-hours of red teaming to precede the launch.

The day before, SpaceXAI announced its new flagship model, Grok 4.5, which it describes as its “most advanced model for coding, agentic work, and knowledge-based tasks.”

An interesting note about Grok 4.5 is that it was built alongside Cursor, an AI coding editor.

This collaboration shines through in the results: according to SpaceXAI, Grok 4.5 bested Anthropic’s Opus 4.8 and Claude Fable in first place on SWE Marathon, a long-horizon benchmark for software engineering, and competed well on Terminal Bench 2.1 and the DeepSWE suite.

Microsoft Azure Unveils World's First NVIDIA GB300 NVL72 Supercomputing  Cluster for OpenAI | NVIDIA Blog

Grok 4.5 was trained on tens of thousands of NVIDIA GB300 GPUs with thousands of hours of fine-tuning on a curated dataset and reinforced learning across hundreds of thousands of multi-step software engineering tasks, resulting, according to SpaceXAI, in “per-token intelligence” as Grok 4.5 “used approximately 4.2x less output tokens to solve SWE Bench Pro tasks compared to Opus 4.8 in max mode.”

To make its entry market, it’s priced aggressively at just $2/million input tokens and $6/million output tokens and will be available free of charge for a limited time on Cursor and Grok Build, and included on all Cursor plans. EU customers can expect availability mid-July; further validation of more significant claims awaits external benchmarks.

Meta jumped into the race with Muse Image, the first media generation model out of its Superintelligence Labs. It’s currently available on the Meta AI app, meta.ai, US Instagram Stories, and in select countries via WhatsApp, with Facebook support coming soon. Beyond its visual output – where it ranks second to OpenAI’s GPT Image 2 on Arena’s text-to-image and image editing leaderboards – Muse Image’s most unique approach to media generation is its agency.

It’s an agent that first researches facts on the real world to ensure its image generation is accurate before coding and executing code to generate precise charts and QR codes.

The agent also self-critiques its work during the generation process and corrects any mistakes naturally.

Meta noted the model’s ability to self-correct stems from its reward function during reinforced learning.

The agent’s images have an imperceptible “Content Seal” watermark, meaning the image won’t alter during manipulation such as cropping or compression and can be detected with an accompanying viewer. Meta is also working on Muse Video, its text-to-video generation model, which is in preview and ranks third in Arena’s text-to-video leaderboards, with a general creator release coming soon. Furthermore, the company released Muse Spark 1.1 and is offering a public preview of the Meta Model API – for the first time offering developers the ability to build directly on top of its Muse model family.

When the technical charts and graphs are stripped away, there’s a single clear signal emerging from these three companies. Every player believes agents, not chatbots, are the future product; every player believes token efficiency is the key competitive advantage, and everyone believes product distribution will ultimately define the winners. Whether that distribution comes in the form of Cursor for SpaceX AI, or Instagram and WhatsApp for Meta, or the enterprise tech stack for OpenAI, distribution will determine it all. And in the past week, prices have gone down, speeds have gone up, and the time between frontier model releases has shrunk from months to days.

Add SpaceRock as a preferred source on Google