All posts tagged: model

Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

Meta says Muse Spark 1.3 has frontier performance — but its best results come from a model developers can’t broadly use yet

Meta’s newest AI model Muse Spark 1.3, unveiled yesterday, is faster and more performant on third-party benchmarks than its predecessor — with a caveat. “Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter,” Meta co-founder and CEO Mark Zuckerberg wrote on X, calling it Meta’s “biggest jump” yet in coding and agentic work. There is substance behind both parts of that claim. Muse Spark 1.3 makes significant gains over last month’s 1.2 release, particularly on long-running agent tasks. The version developers can access now is also one of the strongest price-performance offerings near the top of independent model rankings. Meta’s strongest Muse Spark 1.3 benchmark results come from its max reasoning configuration. Meta says that version is still completing additional safety testing and will arrive “shortly”; the third-party benchmarking firm Artificial Analysis says it evaluated max in a limited partner preview, and currently lists no API provider at all for the configuration. The version broadly rolling out this week through its Muse Code harness and the Meta Model API …

Google’s latest AI weather model gives you no excuse to forget your umbrella

Google’s latest AI weather model gives you no excuse to forget your umbrella

Scientists at Google Deepmind and Google Research released a new artificial intelligence model for weather forecasting today that sees our changing atmosphere more clearly and predicts its behavior more often. WeatherNext 3 is the latest wave of a sea change in meteorology brought out by deep learning techniques, and Google says it will start feeding into weather information users see in search, Google Maps, and Gemini, as well as being available to users and researchers on Google’s cloud platforms. “This is going to be the first time that some of the core variables feed and power a lot of the Google products,” Samier Merchant, a Google senior staff engineer, told TechCrunch. The new model has already proven to be the most accurate among leading contenders tested on Operational WeatherBench, a utility for comparing AI forecasts built by the startup Brightband. It looks at metrics like temperature, windspeed, and humidity. As well as beating out other deep-learning models built by Google, Microsoft, Nvidia, and the European Center for Medium-Range Weather Forecasting, it also beats traditional forecasts …

Open AI’s Astra model is on the way—and very good at breaking into computer systems

Open AI’s Astra model is on the way—and very good at breaking into computer systems

OpenAI shared new details on its forthcoming Astra model, which the company said is the first large language model to meet its “critical cybersecurity threshold,” in preparation for its imminent release. “We plan to make Astra available soon,” OpenAI’s blog post reads, “but access to its most advanced cybersecurity capabilities will be more limited.” The frontier lab determined that Astra is capable of finding unknown security flaws in computer systems, and exploiting them without a person’s guidance. That’s similar to the concerns Anthropic raised about its Mythos model earlier this year, and OpenAI is taking comparable precautions as it prepares to roll out the Astra. Without any third-party confirmation, it is difficult to evaluate OpenAI’s claims about safety or preparedness. The company said it would preview the model with a group of testers, but did not say who they were or how they would be chosen. It’s not clear if OpenAI is working with the US government to evaluate the model ahead of release. OpenAI noted that Astra scored a perfect score on ExploitBench, an …

OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

OpenAI announced Tuesday that its forthcoming AI model, Astra, is its first to reach the company’s threshold for what it calls “critical” cyber capabilities. OpenAI says it plans to publicly release a version of Astra “soon,” but will make the model’s advanced cyber capabilities available only to select partners in its Daybreak Blue early-access program at launch. In a briefing with reporters, OpenAI safety and security leaders said the company has concluded that Astra reaches the critical cybersecurity capabilities outlined in its preparedness framework, which sets thresholds and protocols for when its AI models pose new levels of risk. The company says an AI model has reached its critical cyber threshold when it can independently find and exploit previously unknown vulnerabilities in real-world software. OpenAI leaders said the company has followed its procedure for this situation, which is to halt further development until appropriate safeguards and security measures can be implemented. OpenAI previously said that it paused some training workloads related to the development of Astra and a future AI model for several weeks. Executives …

Kalshi Ninth Circuit ruling threatens federal prediction market model

Kalshi Ninth Circuit ruling threatens federal prediction market model

The Ninth Circuit has called Kalshi’s contracts what its advertising already did: sports betting.  Kalshi’s marketing department wrote the opening line of its own legal defeat. “KalshiEX, LLC advertises itself as ‘the first app for legal sports betting in all 50 states’,” Judge Ryan Nelson wrote in the first sentence of the Ninth Circuit’s 50-page opinion. It appears to be a rough way to learn that ad copy can become evidence. Kalshi has spent much of its legal battle arguing that its sports contracts are not sports bets but financial derivatives traded on a federally regulated exchange. The court noticed the slight branding inconsistency. A unanimous three-judge panel ruled on Friday that Kalshi’s sports event contracts are, in substance, sports bets. Because those bets are not “swaps” within the meaning of the Commodity Exchange Act, they do not fall within the Commodity Futures Trading Commission’s exclusive jurisdiction. Nevada can therefore apply its gaming laws. Calling a wager an event contract, the court effectively concluded, does not make it any less of a wager. Or, as …

ProCook outdoor pizza oven review: a budget-friendly model that brings the pizzeria home | Pizza

ProCook outdoor pizza oven review: a budget-friendly model that brings the pizzeria home | Pizza

The market for pizza ovens tends to be dominated by a handful of specialist brands, so it can feel like a gamble to buy from a lesser-known manufacturer or a retailer that offers a wider range of goods. The ProCook outdoor pizza oven is no gamble, though. It’s every inch a capable, compact and versatile pizza oven, despite being just one product in an offering that includes pretty much everything from storage jars to squeegees. This outdoor pizza oven conforms to ProCook’s usual direct-from-manufacturer pricing. Coming in at a comparatively reasonable RRP of £249, the oven punches far above its weight in terms of functionality. As well as foldable legs for storage, it has room for 12in pizzas and a manual rotating stone for more even cooking. Turn the stone as the pizza bakes for an evenly blistered crust or simply adjust it as needed. At 12kg, it’s also light enough to take to a friend’s house or on a camping trip. View at ProCook How I tested Test slice: one of the veggie pizzas …

Physics-trained AI reveals how Earth’s deep interior changed over time

Physics-trained AI reveals how Earth’s deep interior changed over time

Earth’s mantle moves only a few centimeters each year, but that slow motion helps drive plate tectonics, earthquakes and volcanic activity. A physics-guided AI model reconstructed hidden mantle flow in a controlled simulation using limited surface motion and present-day temperature clues. The work is an early test, not a real-Earth reconstruction, but it could help scientists study how the deep planet changed over time. Deep inside Earth, rock moves at the pace of growing fingernails. That slow circulation shapes plates, feeds volcanoes and influences earthquakes, yet much of its past remains hidden from direct view. The mantle, a rocky layer that accounts for more than 80 percent of Earth’s volume, circulates only a few centimeters per year. Over geological time, that movement helps drive plate tectonics and major surface activity. But scientists cannot simply watch the deep mantle move. They must rely on surface geology and geophysical imaging, including seismic observations, to infer what lies below. A new study by the University of Tsukuba tested whether artificial intelligence could help reconstruct that missing history. The …

Meta researchers taught an 8B AI model to match Claude Opus 4.5 — without the frontier price tag

Meta researchers taught an 8B AI model to match Claude Opus 4.5 — without the frontier price tag

Consider an AI agent tasked with a complex enterprise workflow like migrating massive batches of customer records from a legacy CRM to a cloud database. The agent cannot rely solely on its internal context window for a job spanning hours and depends on the runtime layer, aka the harness. This harness provides execution feedback, like server logs, to help the agent maintain an accurate understanding of dynamic API connections. It also provides state trackers and control-flow mechanisms to manage completed and pending subgoals, ensuring the agent doesn’t skip or duplicate data batches. When unexpected errors occur, such as a database rejecting a batch due to strict API rate limits, the harness provides tools and instructions to help the agent recover. The main way to tell an agent how and when to use its tools is to have a human developer write a set of rules and instructions telling it what to do step-by-step. For example, a developer might instruct the agent to always search the company wiki before writing an email. Because the agent is …

The Powerful Stealth AI Model ‘Ox Alpha’ Is Now GLM-5.3-Flash, and You Can Use It

The Powerful Stealth AI Model ‘Ox Alpha’ Is Now GLM-5.3-Flash, and You Can Use It

The anonymous Ox Alpha AI model that popped up on OpenRouter, the unified AI platform, last week has been claimed by the Chinese company Z.ai. Now officially known as GLM-5.3-Flash, the model is both powerful and cheap, challenging the capabilities and cost of popular models from Anthropic and OpenAI.  When it was first teased, Ox Alpha became the most popular model for the week on OpenRouter and all of its traffic was served through Chinese-made AI chips. After it was confirmed to be the company behind the new model, Z.ai’s stock price soared, Bloomberg reports.  Ox Alpha was described as a “frontier model built for efficient coding, sustained agentic work, and real-world production use,” according to an OpenRouter post on X. In a follow-up post, the platform confirmed that the provider wouldn’t train on your prompts.  Z.ai’s GLM family are open-weight models, where core components are publicly released, allowing anyone to download and use them. (That’s different from open-source AI models, which make all pieces of the model available, including the source code, weights, and …