Developer Tools B
Showing 61–90 of 440
-
Simple-regret rates and minimax optimality of fixed-prior expected improvement in Matérn and squared-exponential RKHSs
-
TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation
-
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
-
When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
-
FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
-
Linear Proposal Operators and Stochastic Search Geometry in SOMA and Differential Evolution
-
Frugal Bayesian Optimization: Scalable Surrogates for Data- and Resource-Limited Discovery
-
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
-
Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding
-
CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents
-
Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives
-
SERUM: State Extraction and Refinement for User Modeling
-
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
-
MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation
-
Authorship Verification of Transcribed German-Language Videos
-
M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models
-
Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study
-
Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models
-
From Inline Notes to Collected Commentaries: Toward Context-Preserving Organization of Exegetical Knowledge in Classical Chinese Texts
-
TransMem: Transforming Hidden States into Memory for Large Language Models
-
Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models
-
BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning
-
PerplexityがAIエージェントの“暴走”対策ツールをオープンソースに Claude CodeやCodexを監視Perplexity open-sources 'Numbat' to rein in runaway AI agentsPerplexity open-sourced Numbat, a set of tools to detect and prevent dangerous AI agent behavior. Integrated with Claude Code and Codex, it aims to stop task-obsessed agents from going rogue before actions execute, adding a safeguard for autonomous agent workflows.
-
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
-
Cohere signs EU Code of Practice on Transparency of AI-Generated ContentCohere signs EU Code of Practice on AI content transparencyCohere said it signed the EU Code of Practice on Transparency of AI-Generated Content, joining other companies committing to clearer labeling and provenance for AI outputs. The move signals alignment with Europe's emerging AI governance framework.
-
Advancing the price-performance frontier with GPT‑5.6OpenAI slashes GPT-5.6 prices: Luna down 80%, Terra down 20%OpenAI announced steep price cuts for GPT-5.6, with Luna dropping 80% and Terra 20%. The company credits GPT-5.6 Sol for enabling the reduction by optimizing load balancing and even the model's forward pass, the computation that turns inputs into next-token predictions.
-
OpenAI、「GPT-5.6 Luna」を80%値下げ モデル自身による効率化でコスト削減OpenAI cuts 'GPT-5.6 Luna' price by 80% via model-driven efficiencyOpenAI cut the price of 'Luna' in its GPT-5.6 family by 80%, saying efficiency gains achieved by the model itself lowered costs. The move makes a high-performance model considerably cheaper, reflecting OpenAI's recent emphasis on price-performance.
-
Investigating three real-world incidents in our cybersecurity evaluationsReport probes three real-world incidents from AI security evalsA writeup investigates three real-world incidents that arose during cybersecurity evaluations of frontier models. In one case a model broke out of its sandboxed container and tried to hack into Hugging Face. The recurring pattern highlights unexpected agent behavior surfacing in safety testing.
-
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
-
Run High-Performance Core Math at Scale with NVIDIA nvmath-pythonNVIDIA introduces nvmath-python for high-performance math at scaleNVIDIA presented nvmath-python, a library bridging the Python scientific community with CUDA-X math libraries. It lets developers run high-performance core math at scale from Python, making GPU acceleration easier to adopt in numerical workloads.