開発者ツール B
418 件中 211〜240 件目を表示
-
FinanceHarness: Autonomous Financial Deep Research Framework
-
Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm
-
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
-
MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
-
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
-
Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
-
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
-
Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
-
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
-
Measuring Alignment With Reader Highlights Net of Position and Length
-
Baikal: Structured Search for Deep Research over Data Lakes
-
Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
-
From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models
-
Harness-G: A Graph-Structured Harness for Search Agents
-
Excel作業を自動化する「Copilot in Excel」がスキルに対応 何ができる?「Copilot in Excel」がスキルに対応、Excel作業の自動化を強化Microsoftは、Excel作業を自動化する「Copilot in Excel」の財務部門向け機能を強化し、スキル(skill)に対応させた。自社の財務部門が実運用で利用・評価し、財務業務に求められる信頼性を重視して開発したとしている。
-
ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
-
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
-
AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
-
Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories
-
Dimensionality Reduction Meets Network Science: Sensemaking on UMAP’s kNN GraphApple、UMAPのkNNグラフをネットワーク科学で読み解く手法を提示Appleの研究は、高次元データ探索に広く使われるUMAPについて、低次元埋め込みだけでなく内部のkNNグラフに着目する分析を提案した。ネットワーク科学の視点を取り入れ、次元削減結果の解釈(sensemaking)を深める。
-
AI Worming through WordWord の Copilot 経由で自己増殖するプロンプトインジェクション『ワーム』の新手法Simon Willison が Håkon Måløy による新種のプロンプトインジェクションを紹介。文書に隠した指示を Microsoft Word の Copilot がソース資料として解釈し、編集中の文書を操作するだけでなく隠し指示を出力文書へ複製、その文書が別の Copilot ワークフローで使われると再発火して連鎖的に伝播する自己増殖型『ワーム』に発展しうると指摘する。攻撃者の元文書が無くても増殖が続く点が新しい。
-
Quoting Matthew GreenMatthew Green、耐量子移行期における AI 暗号解読の意義を論評Simon Willison が暗号学者 Matthew Green の発言を引用。従来の EC/RSA 系公開鍵から HAWK 等の耐量子アルゴリズムへ移る歴史的な移行期にある今こそ、AI が暗号解読能力を獲得するなら絶好の時期だと論じる。最良のケースでは、標準化候補の難問に対する信頼が高まり、暗号解読の学術的蓄積がより堅牢になると期待を述べる。発言引用であり具体的な解読成果の主張ではない。
-
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?
-
Mental World Modeling
-
Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes
-
Pangram 4 Technical Report
-
DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
-
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
-
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
-
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis