AI Models
LLM and model releases, benchmarks, open-weight/open-source models, local AI and model explainers.
AudioShake Launches The Refinery, Speaker-Separated, Quality-Scored Audio for AI Training, by Its Own Account
AudioShake says The Refinery turns mixed recordings into speaker-separated, quality-scored training data. Every figure is the company's own; no independent test.
Top stories
Latest
Two Open Models From 8 October: a Fast Coding Agent and a Gene Finder
JetBrains has released Mellum2.1, a small open model for coding agents, and Hugging Face has released Carbon-A, a 1.2B-parameter model that predicts protein-coding regions in DNA, plus a database of 566 million predicted loci. What each does, and which numbers are the developers' own.
AI Research Tools: TasteVal Measures “Research Taste”, Karotte Targets Reward Hacking and Zephon Fixes Data Order
A new benchmark, TasteVal, says the best AI model now reaches expert-level results on eight AI research tasks with 2.3 times less compute, though its authors list reasons that may overstate progress. Preference Model has open-sourced Karotte, a framework for training environments that are harder to game, and DatologyAI has released Zephon, a data loader that keeps experiments repeatable when the GPU count changes. What each claims, and who measured it.
AI Still Misreads Pictures, and Reasoning May Hinge on One Word: Two New Studies
Scale AI and Elorian's Humanity's Sixth Sense benchmark finds people score 93.1% on everyday visual questions while the best model, GPT-6 Astra, scores 53.6%, mostly because models misread the picture. A separate preprint, not yet peer reviewed, says base models can reason almost as well as RL-trained ones when their answer starts with the right word.
OpenAI’s Decisions API and Perplexity’s Open Decider: AI That Answers With a Probability, Not an Essay
On 6 October OpenAI launched a Decisions API that returns yes/no probabilities, picks and rubric scores from GPT-6 Luna, and cut its API usage tiers from five to three. Perplexity released pplx-decider-v1.1-27b, an Apache 2.0 open-weight decision model. What decision models are, what each costs, and which scores are the vendors' own.
Small Open Models That Answer in One Pass: Liquid AI’s d1 and Perplexity’s pplx-embed-v2-late
Liquid AI open-sourced two "decision models" that give a structured answer in a single pass, in about 8 ms on an RTX 4090 by Liquid's figures. Perplexity open-sourced two search models that read text, images and PDF pages without OCR, where a small model can query an index built by a big one. What each does, in plain words.
Claude Haiku 5.5: Anthropic’s Small Model Gets Much Cheaper and Much Better at Using a Computer
Anthropic's Claude Haiku 5.5 costs around 75% less to run than Haiku 4.5, by Anthropic's estimate, and jumps from 15.7% to 72.4% on Anthropic's OSWorld 2.1 computer-use test. Sonnet 5.5 cache reads are also halved. What is confirmed, what is company benchmark, and who it is for.
GPT-6 Reaches Every ChatGPT Tier: Luna for Free Users, Sol for Paid
OpenAI is rolling GPT-6 into ChatGPT's Chat tab for every plan: GPT-6 Sol for paid tiers from 7 October, GPT-6 Luna for Free and Go from 8 October, plus "Intelligent UI" answers with charts, forms and mini tools. What changed, which Sol this is, and what OpenAI's own safety card flags.
LoGRA Explained: How NVIDIA Cuts RL Training Memory with Low-Rank Gradient Sketches
NVIDIA’s LoGRA paper shows low-rank gradient sketches cutting average RL training memory by up to 45.7% on tested Qwen reasoning setups—and making 27B RL runnable on one 8×H100 node where dense Adam OOMs. What the method is, what Table 1 shows, and what it does not prove.
RAG vs Jev+RAG Explained: What the Judgment Gate Adds
Avi Chawla’s 7 October explainer contrasts standard RAG with a Jev gate between retrieval and generation. What hybrid search already does, what typed probabilities add, why thresholds stay in code, and what this architecture does not replace.
OpenAI’s 372 math and CS results: Lean helps, independent checks pending
Scientific American: OpenAI released 372 math/CS results from an unreleased model on 6 October 2026. Many are Lean-verified; independent check pending.
Latest on TSN
Free TSN tools: crypto calculator, Flux dashboard and more.
