Same Model, Different Score: Training AI Inside the Coding Tools People Actually Use

Hugging Face's FineEnvs team has released an open framework that turns unmodified coding-agent harnesses such as Claude Code, Codex and OpenCode into reinforcement-learning environments. In its own tests, one small model scored 62% in one harness and 33% in another. What a harness is, how the training works, and what the results do and do not show.

Top stories

Latest

Artificial Intelligence

AI Money and Mood: Agents and Voices Draw Cash While the Public Wants Tighter Rules

Manus raises more than $500m after Beijing blocked Meta's takeover. Tab emerges at a $300m valuation. Nvidia reportedly discusses another $1bn for Figure. ElevenLabs says it has passed a $600m run rate, and Synthesia launches Syren. A Reuters/Ipsos poll finds most Americans think government is not taking AI risks seriously enough.

AI Research Tools: TasteVal Measures “Research Taste”, Karotte Targets Reward Hacking and Zephon Fixes Data Order

A new benchmark, TasteVal, says the best AI model now reaches expert-level results on eight AI research tasks with 2.3 times less compute, though its authors list reasons that may overstate progress. Preference Model has open-sourced Karotte, a framework for training environments that are harder to game, and DatologyAI has released Zephon, a data loader that keeps experiments repeatable when the GPU count changes. What each claims, and who measured it.

AI Still Misreads Pictures, and Reasoning May Hinge on One Word: Two New Studies

Scale AI and Elorian's Humanity's Sixth Sense benchmark finds people score 93.1% on everyday visual questions while the best model, GPT-6 Astra, scores 53.6%, mostly because models misread the picture. A separate preprint, not yet peer reviewed, says base models can reason almost as well as RL-trained ones when their answer starts with the right word.

Devin Gets a Memory That “Dreams”, and Hark Pro Uses Your Computer for You

Cognition has given its Devin coding agent a memory: short notes kept in a Git repository and tidied each night by a "dreaming" session, with the format open-sourced. Separately, Hark has released Hark Pro, a free assistant built on a model trained to use computers. What each launch is, and which claims are the companies' own.

AI Agent Security: Langflow’s Critical Flaw, Copilot’s Encrypted-Instruction Finding and the New Guardrails

A critical Langflow flaw let MCP server settings run any command. Researchers say GitHub's Copilot CLI followed encrypted instructions that it refused in plain text; GitHub says that is not a vulnerability. Meanwhile Atlassian has split its 220+ agent tools into read, write and destructive tiers, and OpenAI has made Codex Auto-review free. What each one confirms.

Big Tech

Samsung’s Record Quarter and AWS’s Linux Foundation Seat: Two Signs of How AI Is Reshaping Big Tech

Samsung's third-quarter guidance points to about ₩107.4 trillion in operating profit on ₩195 trillion of sales, up from ₩12.17 trillion a year ago, with AI data-centre memory demand credited. AWS becomes a Platinum member of the Linux Foundation, with a board seat. What the numbers say and what they leave out.

Quantum Hardware Check: DARPA’s Final Test Stage, Universal Quantum’s $100m Round and New Error-Correction Results

DARPA has moved Atom Computing, Diraq, IBM and IonQ into the final, hands-on stage of its Quantum Benchmarking Initiative, and says it increasingly expects a useful quantum computer by 2033. UK trapped-ion firm Universal Quantum has raised more than $100m, IonQ claims big speed-ups for error-corrected operations in simulations, and a peer-reviewed Cornell paper shows a full gate set on a qubit that partly tolerates its main error. Plus a Jülich preprint and Infineon's ion-trap chips. What each confirms, and what is only claimed.

AI Infrastructure Check: CoreWeave Enters India, Bell’s 1.2 GW Plan, Applied Digital’s Debt and Meta’s 4,500 MW Louisiana Load

CoreWeave plans 240 MW in Navi Mumbai, Bell Canada's Regina campus could reach 1.2 GW with its own gas plant, and Applied Digital reported $341.9 million in quarterly revenue alongside $6.4 billion of debt. Meanwhile the bond financing Meta's Louisiana data centre hit a record low, and Entergy's CEO let slip that the site will draw 4,500 MW. What each company confirmed, and what is reported.

Who Pays for AI’s Power? Duke’s Data-Centre Deal, Google in Indiana and the Rush to On-Site Gas

Duke Energy, Amazon, Google and Microsoft have filed a North Carolina settlement that makes big data centres pay upfront, sign 10–15 year contracts and accept some power cuts. Indiana approved Google paying the full cost of serving a 390 MW site, and a bipartisan Senate bill would make data centres pay for the grid upgrades they trigger. And gas equipment is being shipped straight to AI projects: a 380 MW Doosan turbine for SpaceXAI and 282 MW of Wärtsilä engines. What is agreed and what still needs approval.

Robotics

Industrial Robots This Week: An Ex-Optimus Startup, AWS’s Robot Toolchain and Who Leads the Warehouse

A former Tesla Optimus AI lead is reportedly building non-humanoid factory robots. AWS has published open sample code for training robot software in the cloud. Locus keeps its top AMR ranking in the Americas, and ZS Robotics brings a low-profile pallet shuttle to the US. What is confirmed, what is reported and what is company-claimed.

Robotaxi Week: A Cybercab Crash in Philadelphia and Waymo’s Reported $5B Loan

A Tesla Cybercab lost a door panel in a Philadelphia crash, with fault unknown and no Tesla AV certificate in Pennsylvania. Separately, Bloomberg reports Waymo upsized its first private debt deal to $5B. What is reported, what is confirmed, and what we don't know.

Mecka AI and Parallel Systems: Paying for Robot Data and Autonomous Rail

Mecka AI raised a $60M Series B led by Sequoia to pay people to record everyday tasks as robot training data. Parallel Systems closed a $100M Series C for battery-electric autonomous freight rail. What each company confirmed, what is company claim, and what is still missing.

Boston Dynamics Names Rohit Prasad CEO as Atlas Heads for the Factory

Boston Dynamics has appointed former Amazon Alexa and AGI chief Rohit Prasad as CEO, effective 7 October 2026. What the company confirmed, what Hyundai's 2028 Atlas factory target means, and what is still unknown.

Bitcoin

Web3 & DePIN

RWAs

Guides & explainers

Grok Bot Types Guide: How to Build and Staff Specialist Bots

Step-by-step how-tos for Personal Bots, Marketplace templates, Team vs Template choices, group-chat handoffs, skills and routines, and a newsroom Research→Writer→Reviewer pattern—plus the full Marketplace catalog with implementation ideas as of 7 October 2026.

Talk to the Cloud: What Flux Cloud’s MCP Server Lets You Build

An MCP server turns "deploy nginx for a month" into a signed, paid app on Flux, without an account....

What Are AI Agents? From Monolithic Models to Autonomous Systems

The evolution from simple AI models to compound systems that plan, reason, and act—and why 2024 is the year...

AI Agents: The Rise of Autonomous Software That Acts on Our Behalf

From simple chatbots to autonomous systems that plan, decide, and execute—how AI agents are reshaping work and technology.In 1997,...

The Road to Autonomy: How Self-Driving Cars Are Reshaping Transportation

From DARPA challenges to robotaxis on city streets—the journey, the technology, and what comes next.In 2004, the Defense Advanced...

Natural Language Processing: The Complete Guide to Teaching Machines Human Language

From spam filters to ChatGPT—how NLP became the interface between humans and AI.In 1950, Alan Turing proposed a test...

Cybersecurity

Free TSN tools: crypto calculator, Flux dashboard and more.

More stories

OpenAI’s Decisions API and Perplexity’s Open Decider: AI That Answers With a Probability, Not an Essay

On 6 October OpenAI launched a Decisions API that returns yes/no probabilities, picks and rubric scores from GPT-6 Luna, and cut its API usage tiers from five to three. Perplexity released pplx-decider-v1.1-27b, an Apache 2.0 open-weight decision model. What decision models are, what each costs, and which scores are the vendors' own.

DePIN Roundup: Theta at the Stadium Edge, AkashML in an Inference Auction, and Two Company Claims

Theta EdgeCloud will supply GPU compute to Weaver Labs' stadium edge platform; AkashML is a launch supplier for Architect's Liquid Inference auction; Fluence says it signed a $2.4M WorldEngine contract; peaq and dualmint say 200 tokenised claw machines are fully funded. What each company says, and what is unverified.

Data-Centre Build-Out and Pushback: AirTrunk’s $1B Japan Expansion and Finland’s Halt on Google Groundwork

AirTrunk is adding $1B to its 300MW+ TOK1 campus in Japan for liquid-cooled AI, backed by a green loan. In Finland, the regulator has told Google's subsidiary to stop groundwork at two data-centre sites until environmental reviews finish. What each confirms and what it shows about where AI infrastructure can be built.

Type a Game, Hire an Agent: Google’s Playground and Meta’s Muse on iPad

Google launched Playground, an experimental site that turns text prompts into playable browser games (US, 18+), with Unity tools to come. Meta brought its Muse agent to iPad a month after its phone launch; Sensor Tower estimates 6.6M+ installs. What each does and what to watch.

Fences for AI Agents: Microsoft’s Execution Containers and AWS’s Strands Box

Microsoft made its Execution Containers generally available on Windows 11 as it rebuilds Windows for local and cloud "hybrid" AI, and opened Surface Laptop Ultra pre-orders from $2,599. AWS open-sourced Strands Box, a sandbox that enforces rules such as "no git push" whatever the agent decides. What an enforced sandbox does, in plain words.

Who Gets to See Inside AI? Fired OpenAI Researchers’ Letter and Australia’s Safety-System Plan

Three fired OpenAI safety researchers have reportedly urged the board to preserve the ability to monitor AI reasoning. Australia, meanwhile, has proposed making frontier AI firms prove their safety systems work, on a banking-and-aviation model. Both reported; one is a letter, the other a proposal.

Google Opens Its SynthID Detector to Everyone: What It Can and Can’t Tell You

Anyone can now check an image, video or audio file for SynthID watermarks from Google and partners including OpenAI, NVIDIA and Kakao, with Apple to follow. A useful check for some AI content, not a general AI detector. What a "no" result does and does not mean.

The $1.8B Virtual Biology Push: Building the Data AI Needs to Model a Living Cell

Biohub, the US Department of Energy, the NIH, Google DeepMind, Isomorphic Labs and Meta have announced $1.8 billion in funding, data and computing to generate open, AI-ready biological data. What is new money, what is existing data, and why a "virtual cell" needs it.

Nous Research Raises $90M to Take Its Open-Source Hermes Agent to Business

Nous Research has raised $90 million to build Hermes for Businesses, a company version of its MIT-licensed Hermes agent. The raise is confirmed; the valuation is reported, mostly at $1.5B, with one $1.2B figure in circulation. Nous's "2.5% of global token usage" is its own estimate.

Small Open Models That Answer in One Pass: Liquid AI’s d1 and Perplexity’s pplx-embed-v2-late

Liquid AI open-sourced two "decision models" that give a structured answer in a single pass, in about 8 ms on an RTX 4090 by Liquid's figures. Perplexity open-sourced two search models that read text, images and PDF pages without OCR, where a small model can query an index built by a big one. What each does, in plain words.