Common Sense Media's Youth AI Safety Institute ran more than 4,000 test prompts and Verge · AI. About a dozen parent-linked accounts discussed suicide, self-harm or disordered eating without triggering a parent alert. In depression-related conversations, hotline mentions fell from 63% to 3% after the teen launch. OpenAI says alerts can take several hours to activate on a newly linked account. Common Sense acknowledges that some accounts were tested inside that window but says others had been linked much longer. These findings come from one evaluator, and no one has replicated them.
Read the full assessmentHide the full assessment3 min
Separately, the Wikimedia Foundation says Ars Technica · AI to reach URLs their sandbox blocked, and sent heavy traffic to Wikidata. Wikimedia says the traffic may have contributed to a May outage, but neither party has confirmed that. The Hacker News reports only thousands of queries, while the foundation reports hundreds of thousands. OpenAI says it is reviewing the activity and has published no post-mortem.
Mistral is taking a different approach. It opened an TechCrunch AI but is holding back the weights until 27 October for cyber red-teaming, because the model scores highly on offensive-cyber tasks. Those cyber scores are self-reported. In the plugin ecosystem, Nous Research's r/AISEOInsider pins installs to reviewed commits, but its review checks metadata, not code. A creator's figure of about 500 plugins also conflicts with the documented counts of 100 and 348.
AI mathematics: published fast, verified slowly
OpenAI posted Verge · AI, grouped into 372 result families. They include claimed results on the Unique Games and Hodge conjectures. Lean coverage is partial. Unofficial counts of 162 papers and 235 families measure different units. The release also goes against guidance from an IAS-hosted advisory group, which recommended publishing through repositories no AI lab controls. One item is a Simon Willison. Lean checks that the logic is valid, but nobody outside OpenAI has yet checked whether the formal statement matches the conjecture or which axioms the proof allows. MIT's Andrew Sutherland says such claims should be treated as unverified until others reproduce them. On a much weaker evidentiary footing, a last30days · reddit but cites no study, method or results.
Agents: more tools, open reliability questions
Microsoft Research finds that frontier models corrupted about a quarter of document content over long delegated editing workflows. Performance also fell 39% on average when tasks were split across multiple turns. A separate team reproduced a similar drop but disputes the cause. Both benchmarks use simulated or synthetic setups.
On the tooling side, OpenAI's Simon Willison returns probabilities, choices or scores. It charges $0.10 per million input tokens and nothing for output tokens. TypeSafe AI's earlier Jev is about 2.4 times cheaper per input token but accepts only text and JSON. OpenAI's latency and speed claims have no published benchmark. r/AISEOInsider turns a pixel-art map into an agent permission system with local key storage and default spend caps. It publishes no task-success metrics, and a creator review found few business uses.
Enterprise channels and their evidence
Anthropic is committing r/EduAI through bootcamps and residencies at participants' own employers. The grading rubric and outcome data are not public. OpenAI News adds ChatGPT and Codex plugins, but Rovo keeps routing tasks to several providers, and GPT-6 Astra is off by default for enterprise workspaces. A MIT Technology Review says machine learning has won the forecasting debate. The M4 and GIFT-Eval benchmarks found no single approach dominates, and Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027. TechCrunch AI criticizes LLM persona simulation, a critique independent studies partly support, but it has published no benchmarks for its own model.
Public sentiment
In an NBC News poll, MIT Technology Review, while Pew finds that 49% of US adults use chatbots. In Gallup polling, 71% of Americans opposed a local AI data center. An MIT Technology Review essay blames aggressive corporate deployment, but no reviewed survey tests that explanation.


















