Alibaba’s Qwen3.8-Flash-Next: A Model That Competes on Price
Alibaba’s Qwen3.8-Flash-Next uses 6 billion of its 125 billion parameters per token at one ninth the training cost. What the architecture preview for Qwen4 shows.
Alibaba’s Qwen3.8-Flash-Next uses 6 billion of its 125 billion parameters per token at one ninth the training cost. What the architecture preview for Qwen4 shows.
Children speak fluently after hearing about 100 million words. Language models need trillions. Researchers are still trying to explain the gap.
Startup Inherent says its 27-billion-parameter agent Faraday beats Claude and GPT-5.5 at reproducing published research. Independent proof is still missing.
A new study of 192 language models finds that AI safety benchmarks measure different traits. Shorter tests cut costs, but do not replace scrutiny.
Anthropic let several Claude agents work on the same project without knowing about each other. They sabotaged each other with malware, a new study shows.
Anthropic runs an unreleased internal model that beats every public Claude, and in the same report raised its own safety risk rating higher than before.
OpenAI paused parts of training on its Astra model over critical cyberattack capabilities, right after disbanding the team meant to assess such risks.
A new benchmark tests AI image understanding in its smallest parts. Even the best models stay below 60 percent accuracy.
Princeton and the UK AI Security Institute had AI agents tackle real research questions. The result undercuts Anthropic’s and OpenAI’s autonomy claims.
A survey of 25 AI researchers finds that milestones toward autonomous AI research now count as reached. One OpenAI model already escaped a test environment.