Developer Tools B
Showing 241–270 of 417
-
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark
-
DLAM: Distributional Latent Actions with Temporal Constraints
-
Linguistic Monoculture in LLM-Assisted Language Use
-
Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes
-
AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching
-
Hierarchical Spatio-Temporal Transformer for Coherent Emergency Department Forecasting
-
Detecting seizure onset and offset times using human intelligence: A critical-transitions-based approach
-
Sky sphere representation in language models
-
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
-
MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair
-
Field Codes for Distributed Coupling Samplers and Certified Empirical Transport
-
Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering
-
Single-Beat Cuffless Blood Pressure Estimation Using Ear-PPG and ECG with a Lightweight Hybrid Learning Framework
-
Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise
-
Visual Credit Audit for Multimodal Spatial Reasoning
-
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data
-
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation
-
GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding
-
Mitigating Compounding Error via Video Representation Regularization
-
Lottery Tickets Are Not Deployment Tickets
-
HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models
-
Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making
-
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
-
How enabling two settings tripled our scores on the ARC-AGI-3 benchmarkOpenAI triples ARC-AGI-3 scores by enabling two API settingsOpenAI reported that enabling two API settings tripled GPT-5.6's scores on the ARC-AGI-3 benchmark. By retaining reasoning across calls and tuning configuration, the setup improved both accuracy and efficiency, which the post breaks down in detail.
-
On the robustness of noisy solutions in non-convex neural networks
-
AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents
-
A Compositional Theory of Causally Masked Transformers
-
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception
-
Using large language models to probe the limits of atom-centered structural descriptors
-
OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment