開発者ツール B
430 件中 91〜120 件目を表示
-
Run High-Performance Core Math at Scale with NVIDIA nvmath-pythonNVIDIA、大規模数値計算向けnvmath-pythonを紹介NVIDIAは、Pythonの科学計算コミュニティとCUDA-Xの数値計算ライブラリを橋渡しするnvmath-pythonを紹介した。高性能なコア数値計算をPythonから大規模に実行でき、GPUアクセラレーションの活用を容易にする。
-
TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
-
Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications
-
Science One Framework: A verifiable autonomous research framework via Chain-of-EvidenceGoogle Research、証拠連鎖で検証可能な自律研究基盤Science Oneを発表Google Researchは、科学研究を自律的に進めるフレームワークScience Oneを発表した。推論の各段階を証拠の連鎖(Chain-of-Evidence)として検証可能にすることで、AIによる研究成果の追跡性と信頼性を高めることを狙う。
-
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
-
Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing
-
Self-Supervised Skill Optimization
-
The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation
-
Quoting Bruce SchneierBruce Schneier、「執筆は筋トレ」とAI時代の学びを説くSimon Willison氏がBruce Schneier氏の言葉を引用。学生への作文課題は「成果物」ではなく「筋トレ」であり、書く行為そのもの(考え、構成し、下書きし、推敲する過程)にこそ意味があると論じる。AIに執筆を任せる時代における学習の価値を問う内容。
-
Learning to Trace Seiberg Dualities
-
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
-
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
-
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
-
Inducing language models to assert their own consciousness restores human beliefs and values
-
Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments
-
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
-
$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
-
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
-
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
-
Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs
-
ORCA-bench: How Ready Are Language Model Agents for Oncall?
-
AI systems and the reproduction of (standard) language ideologies in World Englishes
-
Selective Credibility-Limited Belief Update
-
Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
-
Creative Transformation in Literary Texts: Modelling Change Across Representational Levels
-
InfoOps Bench: A live information operations safety benchmark
-
The Role of Causality in Algorithmic Recourse
-
Beyond Sentiment: Structured Information Extraction from Financial News
-
Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation
-
A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks