·
SOURCES
TRENDING · LAST 24H
  • ·vLLM adds Hy4-preview and PLaMo3 speculative decoding support; SGLang enables Qwen 3.8 Flash ngram lookup offloading to SSD
  • ·NVIDIA Vera Rubin NVL72 hits 30x higher throughput per megawatt via AgentX; Samsung integrates MAC units into LPDDR5X-9600 memory
  • ·SKILL.state replaces conversational history with mutable execution state; CritICL uses small model failures for inference-time reasoning scaling
  • ·Lambda secures $1B debt for Nvidia chip inventory; OpenAI ends model access for Cursor following SpaceX acquisition
  • ·Conduct provides runtime governance for MCP tools; only-cli reduces web-browsing token usage by 142x compared to raw HTML
#1[@OpenAI]
·
16h ago
OpenAI to end model access for Cursor following SpaceX acquisition

OpenAI will terminate its partnership with Cursor on November 12 following the editor's acquisition by SpaceX. This move will end Cursor's direct access to OpenAI models.

breakdown →
#2[HUGGINGFACE]
CritICL Framework Uses Small Model Failures for Inference-Time Reasoning

CritICL improves reasoning performance by using failure modes from smaller models in a family as in-context guidance for larger models. This approach enables inference-time scaling without the heavy computational overhead of repeated generation or external verification.

breakdown →
#3[GH]
·
13h ago
vLLM Adds Support for Hy4-preview Model
★ 0 new · 0 total

The vLLM inference engine now supports the Hy4-preview model. This integration enables high-throughput serving of the architecture within the existing vLLM ecosystem.

breakdown →
#4[TECHCRUNCH]
3h ago
Nvidia shifts focus to data center traffic control efficiency

Nvidia is moving beyond pure GPU compute to improve data center efficiency through smarter traffic control mechanisms. This approach prioritizes optimized data movement and system-level orchestration over simply increasing processor cycles.

breakdown →
#5[APPLE_ML]
Quantifying Bayesian inconsistency in large language models

New research introduces an information processing gap metric to measure how much LLM belief updates deviate from ideal Bayesian updates. The method treats LLMs as information processing rules to evaluate their ability to rationally update probabilistic beliefs when presented with new evidence.

breakdown →
#6[arXiv]
SKILL.state Architecture Replaces Conversational History with Mutable Execution State
cs.AI, cs.MA

SKILL.state improves long-horizon agent performance by using a structured, mutable execution state instead of an ever-growing conversation history. This approach discards intermediate reasoning after state updates, reducing cumulative token consumption and preventing context poisoning.

breakdown →
#7[HN]
3h ago
Debian adopts policy for responsible generative AI use
52 pts · 22 comments

The Debian Project has voted to allow the use of large language models for development, maintenance, and documentation. Contributors remain fully responsible for ensuring all AI-assisted code meets existing quality, correctness, and legal compliance standards.

breakdown →
#8[r/LocalLLaMA]
·
10h ago
Qwen 3.8 Flash Next Ngram Lookup Offloading to SSD
184 upvotes · 33 comments

SGLang supports offloading Qwen 3.8 Flash Next ngram lookup tables to SSD with streaming capabilities. This technique aims to reduce memory overhead during inference with minimal performance degradation.

breakdown →
#9[@nvidia]
3h ago
NVIDIA Vera Rubin NVL72 achieves 30x higher throughput per megawatt via AgentX

NVIDIA's Vera Rubin NVL72 architecture demonstrates up to 30x better throughput per megawatt compared to GB300 NVL72 when running the SemiAnalysis AgentX workload. This benchmark focuses on energy-efficient throughput specifically for agentic workflows rather than traditional training metrics.

breakdown →
#10[HUGGINGFACE]
Luce Generates Relightable 3D Assets via Multimodal Gaussian Clouds

Luce uses a voxelized multimodal Gaussian cloud to unify geometry and PBR materials like albedo and surface normals. A rectified-flow transformer generates a material-aware latent space from a single image, enabling integration into standard rendering pipelines.

breakdown →
#11[GH]
18h ago
vLLM Adds Speculative Decoding Support for PLaMo3
★ 0 new · 0 total

The vLLM inference runtime now supports the speculative decoding method for the PLaMo3 model. This integration allows for faster token generation by using a smaller draft model to predict sequences.

breakdown →
#12[TECHCRUNCH]
18h ago
Lambda Secures $1B Debt to Expand Nvidia Chip Inventory

Lambda secured $1B in private debt to acquire additional Nvidia AI chips for leasing to Microsoft. This capital injection highlights the increasing reliance on debt financing to meet the massive hardware requirements of the AI infrastructure boom.

breakdown →
#13[HUGGINGFACE]
Hugging Face Open ASR Leaderboard Expands to Global South Languages

The Open Automatic Speech Recognition (ASR) Leaderboard has added its first language from the Global South. This expansion improves benchmarking visibility for speech models in previously underrepresented linguistic regions.

breakdown →
#14[arXiv]
Evaluating LLM Summaries for Accuracy in Cancer Care
cs.CL, cs.AI

An evaluation of AI-generated summaries for cancer patients identified risks regarding clinical accuracy and omissions. Researchers used a combination of oncology clinician assessments and LLM-as-a-judge to iteratively improve prompt grounding and safety guardrails.

breakdown →
#15[HN]
22h ago
Conduct provides runtime governance for LLM and MCP tool calls
4 pts · 0 comments

Conduct offers a policy engine and LLM proxy to enforce block, warn, audit, or inject actions across AI agents and shell tools. It distinguishes itself from observability tools by providing fail-closed runtime governance with signed configurations and SHA-256 hash-chained audit logs.

breakdown →
#16[r/ClaudeCode]
19h ago
Jean-Claude Proxy Bypasses Claude Enterprise Managed Settings
343 upvotes · 37 comments

Jean-Claude is a Node.js MITM proxy designed to intercept Claude Code requests for enterprise-managed settings. It allows users to bypass administrative restrictions, such as disabled auto-mode or mandatory tool permissions, by serving custom configurations during the session.

breakdown →
#17[@emollick]
1h ago
Gemini-powered Co-Scientist extension accelerates multi-disciplinary scientific discovery

Google is extending the Co-Scientist framework using Gemini to automate workflows in materials science, biology, and computer science. The system aims to bridge existing gaps in autonomous AI scientific reasoning through human-AI collaboration across diverse research domains.

breakdown →
#18[HUGGINGFACE]
UrbanGround Benchmark for Multimodal City Navigation Agents

UrbanGround provides a 3D geospatial sandbox based on Hong Kong to test how MLLM agents translate local street-view perception into long-term spatial navigation. The framework enables closed-loop, first-person interaction within a physically constrained digital twin.

breakdown →
#19[WIRED]
22h ago
AI performance in medical domains challenges human practitioner capabilities

Recent findings suggest AI systems are outperforming human doctors in specific diagnostic and clinical tasks. The rising efficacy of these models poses significant professional challenges to the medical community.

breakdown →
#20[arXiv]
Freshness-Bounded Shield Mitigates TOCTOU Hazards in LLM Guardrails
cs.AI

LLM guardrails in self-adaptive systems suffer from verdict staleness, with error rates reaching 48.4% after just eight simulator steps. The Freshness-Bounded Shield (FBS) estimates the validity horizon of an approval based on feature volatility to prevent unsafe executions.

breakdown →
#21[HN]
11h ago
Samsung LPDDR5X-PIM Integrates MAC Units into DRAM Banks
4 pts · 0 comments

Samsung implements MAC units directly within LPDDR5X-9600 memory banks to exploit high internal bandwidth and reduce DRAM-to-core latency. The architecture maintains compatibility with standard memory controllers while enabling in-memory computation across 16 banks.

breakdown →
#22[r/ClaudeAI]
only-cli reduces web-browsing token usage by 142x
151 upvotes · 33 comments

The only-cli tool converts websites into compact, agent-ready formats, consuming 142x fewer tokens than raw HTML. It claims to be significantly cheaper than Claude Code's WebSearch and successfully unblocks sites like Reddit and LinkedIn.

breakdown →
#23[@pk_iv]
23h ago
OpenAI Research Claims 80% US GDP Potential for Agentic Automation

OpenAI research suggests agentic systems could automate 80% of US GDP, yet deployment remains constrained by hardware bottlenecks. Discussion focuses on the transition from model development to domain-specific agent labs and the impact of chip and memory shortages.

breakdown →
#24[TECHCRUNCH]
22h ago
Open-weight model startups emerge as primary acquisition targets

Capital flows into open-weight AI companies are driving a trend of acquisitions by larger industry players. This shift highlights a market strategy focused on acquiring high-quality model weights and talent through the open-source ecosystem.

breakdown →
#25[arXiv]
Survival-Guided Length Control Speeds Diffusion LM Inference 7x
cs.CL

A training-free, plug-in length predictor treats end-of-sequence token selection as a discrete-time survival problem. This method accelerates inference by up to 7x across reasoning and code benchmarks while maintaining task accuracy.

breakdown →
9 sources · live pipeline status →↓ mac menu bar app (apple silicon)
·