·
SOURCES
#1[HN]
·
2d ago
ChatGPT Health reaches general availability for consumer medical data analysis
4 pts · 1 comments

OpenAI has transitioned ChatGPT Health to general availability, allowing users to upload medical records and test results for direct analysis. This rollout moves health-related querying from manual file uploads into a specialized interface, potentially disintermediating niche health-tech startups.

breakdown →
#2[7MIN.AI]
·
2d ago
OpenAI Secures 3.2GW Power Deal for Georgia Data Center Campus

OpenAI's Project Camellia involves a massive data center development in Effingham County, Georgia. The project includes a power agreement for 3.2 gigawatts of capacity spanning through 2032.

breakdown →
#3[@claudeai]
·
1d ago
Claude Opus 5 released at half the price of Fable 5

Claude Opus 5 provides frontier-level intelligence comparable to Fable 5 but at 50% of the cost. The model is designed for proactive and thoughtful task execution.

breakdown →
#4[THEVERGE]
·
4d ago
Google releases Gemini 3.5 Flash Cyber for security patching

Google launched Gemini 3.5 Flash Cyber to identify and patch security vulnerabilities at a lower cost than models like Anthropic's Mythos. The model leverages the Flash architecture to provide high-speed security analysis for enterprise environments.

breakdown →
#5[HUGGINGFACE]
·
2d ago
Nunchaku 4-bit Quantized Diffusion Inference Integrated into Diffusers Library

The Diffusers library now supports Nunchaku 4-bit quantization for diffusion models. This integration enables higher-speed inference and reduced VRAM requirements for running large-scale diffusion workloads.

breakdown →
#6[HUGGINGFACE]
SkewAdam Reduces MoE Training Memory by 97% via Tiered Allocation

SkewAdam optimizes Mixture-of-Experts memory by applying differentiated optimizer states to the backbone, experts, and router. For a 6.78B parameter model, it reduces optimizer state from 50.6 GB to 1.29 GB, cutting peak training memory from 81.4 GB to 31.3 GB.

breakdown →
#7[GH]
·
2d ago
Alibaba open-sources hybrid LLM and deterministic code review tool
★ 265 new · 11,194 total

Alibaba released open-code-review, a tool combining deterministic pipelines with LLM agents to provide line-level comments. It includes fine-tuned rulesets for detecting NPE, thread-safety, XSS, and SQL injection, with compatibility for OpenAI and Anthropic APIs.

breakdown →
#8[arXiv]
·
HPD-Parsing Enables Hierarchical Parallel Decoding for Document Parsing
cs.CL

HPD-Parsing implements a hierarchical parallel decoding paradigm for Vision-Language Models to overcome sequential bottlenecks in document parsing. A layout branch organizes structure while concurrent branches decode block-level content using progressive multi-token prediction.

breakdown →
#9[r/LocalLLaMA]
·
2h ago
Llama.cpp Adds Full MCP Support for Agentic Workflows
125 upvotes · 28 comments

Llama.cpp now integrates the Model Context Protocol (MCP), enabling full support for stdio and HTTP servers. Users can configure MCP servers via JSON or command-line arguments to transform the WebUI into an agentic chat interface, such as a local coder using the Serena server.

breakdown →
#10[HN]
·
4d ago
Kimi K3 achieves 93% accuracy via routing with Fable 5
24 pts · 2 comments

Kimi K3 achieves frontier-level performance by routing agentic tasks between the open K3 model and the closed Fable 5 model. This hybrid approach maintains 93% accuracy on 1,000 tasks while offering up to 50x cost reductions via Fireworks hosting.

breakdown →
#11[TLDR]
·
1d ago
Stripe Negotiating $10 Billion Acquisition of OpenRouter

Stripe is in discussions to acquire OpenRouter, a model routing platform for developers. The deal is reportedly valued at approximately $10 billion.

breakdown →
#12[@OpenAI]
·
3d ago
OpenAI Presence Enables Enterprise Deployment of Voice and Chat Agents

OpenAI Presence provides enterprise customers with tools to deploy voice and chat agents that interact with internal company systems and execute approved actions. The service is currently available through a limited general availability program for eligible organizations.

breakdown →
#13[THEVERGE]
·
5d ago
Moonshot and Alibaba release low-cost competitive AI models

Chinese AI firms Moonshot and Alibaba have released new models targeting parity with OpenAI and Anthropic performance levels. These releases prioritize significant cost reductions compared to leading US-based frontier models.

breakdown →
#14[OPENAI]
·
3d ago
OpenAI partners with US Department of Energy for scientific discovery

OpenAI is collaborating with the U.S. Department of Energy and national laboratories to apply frontier models to scientific research. The initiative aims to accelerate discovery across various national science domains.

breakdown →
#15[HUGGINGFACE]
SANA-Video 2.0 Hybrid Linear Attention for 720p Video Generation

SANA-Video 2.0 uses a hybrid architecture with gated linear attention and periodic gated-softmax anchors at a 3:1 ratio to achieve O(N) scaling. The 5B and 14B parameter models generate 720p video on a single GPU by using Block Attention Residuals to propagate feature updates.

breakdown →
#16[GH]
6d ago
WrenAI Open-Source Generative BI via Text-to-SQL
★ 96 new · 16,088 total

WrenAI implements a governed text-to-SQL engine using an open context layer to generate dashboards and charts from natural language. It supports over 20 data sources, including Snowflake, BigQuery, and PostgreSQL.

breakdown →
#17[arXiv]
·
CLARK: Closed-Loop Reasoning via Knowledge Graphs and Markov Logic Networks
cs.AI

Integrates knowledge graphs with symbolic rule mining and probabilistic reasoning using the LP-MLN formalism. The framework iteratively enriches graph structures with candidate rules that are calibrated through probabilistic weight learning to handle uncertainty.

breakdown →
#18[r/LocalLLaMA]
·
2d ago
W8A8 Kernels Enable 1.4x Speedup on Apple M5 Silicon
105 upvotes · 42 comments

Custom W8A8 kernels for Apple M5 hardware achieve 3,029 tps for Gemma4 prefill tasks, a 1.4x improvement over the 2,193 tps baseline. Current inference backends like MLX and Llama.cpp lack support for the M5's native INT8 activation capabilities.

breakdown →
#19[HN]
·
4d ago
Meta AI Models Integrated into Lawrence Berkeley National Laboratory Research
12 pts · 0 comments

Meta's AI models are being deployed to process petabyte-scale datasets generated by the Advanced Light Source X-ray beamlines. The integration aims to manage increasing data volumes from high-resolution facility upgrades that exceed manual scientific analysis capacities.

breakdown →
#20[TLDR DEV]
·
4d ago
Planner-Worker Swarms Improve Software Engineering Performance

Implementing a planner-worker architecture in agent swarms improves software development output quality. This structure allows for cost savings by delegating execution tasks to lower-cost models while maintaining high-level reasoning.

breakdown →
#21[@sundarpichai]
·
3d ago
Alphabet Q2 Results: Gemini APIs reach 22B tokens per minute

Alphabet reported 24% YoY revenue growth with Google Cloud accelerating at 82%. Gemini model APIs are now processing 22 billion tokens per minute, driven largely by Flash model adoption.

breakdown →
#22[WIRED]
·
3d ago
Chinese open-source models challenge Silicon Valley frontier dominance

Chinese labs are releasing open-source models to provide stable, accessible alternatives to restricted US models from Anthropic and OpenAI. These models aim to capture market share through availability and increasing capability.

breakdown →
#23[APPLE_ML]
LEAD method addresses no-recovery bottleneck in long-horizon reasoning

The Lookahead-Enhanced Atomic Decomposition (LEAD) method mitigates error accumulation in long-horizon LLM tasks. By incorporating short-horizon future validation during step decomposition, it prevents irreversible errors caused by non-uniform error distributions in complex algorithmic puzzles.

breakdown →
#24[HUGGINGFACE]
Xiaomi-Robotics-1 VLA Model Trained on 100K Hours of Trajectories

Xiaomi-Robotics-1 is a vision-language-action model trained on 100,000 hours of real-world manipulation data via an auto-labeling pipeline. It enables zero-shot mobile manipulation and rapid adaptation to new tasks with minimal fine-tuning data.

breakdown →
#25[GH]
5d ago
fastmcp provides Pythonic Model Context Protocol implementation
★ 77 new · 26,420 total

fastmcp is a Python library designed for rapid development of Model Context Protocol (MCP) servers and clients. It streamlines the creation of standardized interfaces for AI agents to access external data and tools.

breakdown →
9 sources · live pipeline status →↓ mac menu bar app (apple silicon)
·