Developer Tools B
Showing 91–120 of 430
-
Run High-Performance Core Math at Scale with NVIDIA nvmath-pythonNVIDIA introduces nvmath-python for high-performance math at scaleNVIDIA presented nvmath-python, a library bridging the Python scientific community with CUDA-X math libraries. It lets developers run high-performance core math at scale from Python, making GPU acceleration easier to adopt in numerical workloads.
-
TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
-
Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications
-
Science One Framework: A verifiable autonomous research framework via Chain-of-EvidenceGoogle unveils Science One, a verifiable autonomous research frameworkGoogle Research introduced Science One, a framework for autonomous scientific research that makes each reasoning step verifiable through a Chain-of-Evidence. The design aims to improve the traceability and trustworthiness of AI-driven research results.
-
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
-
Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing
-
Self-Supervised Skill Optimization
-
The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation
-
Quoting Bruce SchneierBruce Schneier: writing assignments are 'gym tasks,' not workSimon Willison quotes Bruce Schneier arguing that student writing assignments are 'gym tasks, not work tasks.' The value lies in the act of writing itself, thinking, outlining, drafting, and editing, rather than the output, a pointed reflection on learning in the age of AI writing tools.
-
Learning to Trace Seiberg Dualities
-
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
-
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
-
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
-
Inducing language models to assert their own consciousness restores human beliefs and values
-
Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments
-
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
-
$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
-
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
-
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
-
Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs
-
ORCA-bench: How Ready Are Language Model Agents for Oncall?
-
AI systems and the reproduction of (standard) language ideologies in World Englishes
-
Selective Credibility-Limited Belief Update
-
Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
-
Creative Transformation in Literary Texts: Modelling Change Across Representational Levels
-
InfoOps Bench: A live information operations safety benchmark
-
The Role of Causality in Algorithmic Recourse
-
Beyond Sentiment: Structured Information Extraction from Financial News
-
Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation
-
A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks