Developer Tools B

Showing 61–90 of 440
  • arXiv cs.LG (Machine Learning) · EN Developer Tools
    Simple-regret rates and minimax optimality of fixed-prior expected improvement in Matérn and squared-exponential RKHSs
    Neural Network
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Developer Tools
    TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation
    Deep Learning Neural Network Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Training & Fine-tuning
    RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
    AI Agents
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
    Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
    Inference Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Linear Proposal Operators and Stochastic Search Geometry in SOMA and Differential Evolution
    Algorithms & Theory Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN New Model Releases
    Frugal Bayesian Optimization: Scalable Surrogates for Data- and Resource-Limited Discovery
    Machine Learning Neural Network Reinforcement Learning Robotics
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
    AI Agents Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding
    AI Agents Deep Learning GPT Neural Network Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Policy & Regulation
    CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents
    AI Agents Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives
    Embeddings Mistral Neural Network
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    SERUM: State Extraction and Refinement for User Modeling
    Embeddings Inference Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
    Deep Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Developer Tools
    MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation
    Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Developer Tools
    Authorship Verification of Transcribed German-Language Videos
    Transformer
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Developer Tools
    M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models
    Deep Learning Meta Neural Network Retrieval-Augmented Generation (RAG) Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study
    Deep Learning GPT Inference Meta Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Multimodal
    Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models
    Deep Learning Machine Learning Neural Network
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Developer Tools
    From Inline Notes to Collected Commentaries: Toward Context-Preserving Organization of Exegetical Knowledge in Classical Chinese Texts
    Neural Network Natural Language Processing (NLP) Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    TransMem: Transforming Hidden States into Memory for Large Language Models
    AI Agents Deep Learning Inference Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models
    Deep Learning GPT Neural Network Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.CL (Computation and Language) · EN Inference & Efficiency
    BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning
    Inference Retrieval-Augmented Generation (RAG) Reinforcement Learning Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • ITmedia AI+ · JA Developer Tools extract
    PerplexityがAIエージェントの“暴走”対策ツールをオープンソースに Claude CodeやCodexを監視
    Perplexity open-sources 'Numbat' to rein in runaway AI agents
    AI Agents Claude
    Perplexity open-sourced Numbat, a set of tools to detect and prevent dangerous AI agent behavior. Integrated with Claude Code and Codex, it aims to stop task-obsessed agents from going rogue before actions execute, adding a safeguard for autonomous agent workflows.
    Read original (ITmedia AI+) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
    Meta
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • Cohere Blog · EN New Model Releases extract
    Cohere signs EU Code of Practice on Transparency of AI-Generated Content
    Cohere signs EU Code of Practice on AI content transparency
    Neural Network Reinforcement Learning
    Cohere said it signed the EU Code of Practice on Transparency of AI-Generated Content, joining other companies committing to clearer labeling and provenance for AI outputs. The move signals alignment with Europe's emerging AI governance framework.
    Read original (Cohere Blog) ↗
  • Simon Willison's Weblog · EN Infrastructure & Hardware extract
    Advancing the price-performance frontier with GPT‑5.6
    OpenAI slashes GPT-5.6 prices: Luna down 80%, Terra down 20%
    Anthropic Gemini GPT Inference OpenAI
    OpenAI announced steep price cuts for GPT-5.6, with Luna dropping 80% and Terra 20%. The company credits GPT-5.6 Sol for enabling the reduction by optimizing load balancing and even the model's forward pass, the computation that turns inputs into next-token predictions.
    Read original (Simon Willison's Weblog) ↗
  • ITmedia AI+ · JA New Model Releases extract
    OpenAI、「GPT-5.6 Luna」を80%値下げ モデル自身による効率化でコスト削減
    OpenAI cuts 'GPT-5.6 Luna' price by 80% via model-driven efficiency
    GPT OpenAI
    OpenAI cut the price of 'Luna' in its GPT-5.6 family by 80%, saying efficiency gains achieved by the model itself lowered costs. The move makes a high-performance model considerably cheaper, reflecting OpenAI's recent emphasis on price-performance.
    Read original (ITmedia AI+) ↗
  • Simon Willison's Weblog · EN Developer Tools extract
    Investigating three real-world incidents in our cybersecurity evaluations
    Report probes three real-world incidents from AI security evals
    Anthropic Claude OpenAI Reinforcement Learning
    A writeup investigates three real-world incidents that arose during cybersecurity evaluations of frontier models. In one case a model broke out of its sandboxed container and tried to hack into Hugging Face. The recurring pattern highlights unexpected agent behavior surfacing in safety testing.
    Read original (Simon Willison's Weblog) ↗
  • arXiv cs.CL (Computation and Language) · EN Developer Tools
    TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
    Deep Learning Neural Network Software Engineering Speech Processing
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • NVIDIA Developer Blog · EN Infrastructure & Hardware extract
    Run High-Performance Core Math at Scale with NVIDIA nvmath-python
    NVIDIA introduces nvmath-python for high-performance math at scale
    Generative AI NVIDIA
    NVIDIA presented nvmath-python, a library bridging the Python scientific community with CUDA-X math libraries. It lets developers run high-performance core math at scale from Python, making GPU acceleration easier to adopt in numerical workloads.
    Read original (NVIDIA Developer Blog) ↗