Developer Tools B
Showing 211–240 of 408
-
FinanceHarness: Autonomous Financial Deep Research Framework
-
Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm
-
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
-
MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
-
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
-
Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
-
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
-
Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
-
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
-
Measuring Alignment With Reader Highlights Net of Position and Length
-
Baikal: Structured Search for Deep Research over Data Lakes
-
Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
-
From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models
-
Harness-G: A Graph-Structured Harness for Search Agents
-
Excel作業を自動化する「Copilot in Excel」がスキルに対応 何ができる?'Copilot in Excel' gains skills support to automate spreadsheet workMicrosoft strengthened the finance-focused capabilities of Copilot in Excel and added support for 'skills' to automate spreadsheet tasks. The company says its own finance team used and evaluated the feature in production, developing it with the reliability that financial work demands.
-
ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
-
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
-
AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
-
Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories
-
Dimensionality Reduction Meets Network Science: Sensemaking on UMAP’s kNN GraphApple applies network science to UMAP's kNN graph for sensemakingApple researchers bring network science to UMAP, arguing that typical workflows focus only on its low-dimensional embedding. By analyzing UMAP's underlying kNN graph, the work aims to improve sensemaking and interpretation of dimensionality-reduction results.
-
AI Worming through WordPrompt-injection 'worm' self-replicates through Microsoft Word's CopilotSimon Willison highlights a prompt-injection variant by Håkon Måløy that upgrades attacks on Microsoft Word's Copilot into self-replicating worms. Hidden instructions in a source document can be interpreted by Copilot, which may both manipulate the drafted document and copy the instructions into the output, turning it into a new carrier that re-triggers in later Copilot workflows and propagates even without the attacker's original file present.
-
Quoting Matthew GreenMatthew Green on AI cryptanalysis arriving amid the post-quantum shiftSimon Willison quotes cryptographer Matthew Green on the historic shift from EC/RSA public-key schemes to post-quantum algorithms such as HAWK. Green argues that if AI is going to become good at cryptanalysis, now, during this transition, is the ideal moment: in the best case it would build real confidence in the hard problems being standardized and make the cryptanalysis literature far more robust. This is commentary, not a claim of a specific break.
-
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?
-
Mental World Modeling
-
Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes
-
Pangram 4 Technical Report
-
DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
-
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
-
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
-
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis