Safety & Evaluation

A
Showing 31–58 of 58
  • arXiv cs.LG (Machine Learning) · EN New Model Releases
    Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication
    Deep Learning Quantization
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Safety & Evaluation
    Recurrent GraphNeural NetworkswithSet-BasedAggregation
    Neural Network Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN New Model Releases
    Safe Meta-Reinforcement Learning via Information Space Reachability
    AI Agents Meta Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN Safety & Evaluation
    Inoculation Midtraining with Learned Neologisms
    Fine-tuning Reinforcement Learning
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Training & Fine-tuning
    K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
    GPT Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Safety & Evaluation
    CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering
    Retrieval-Augmented Generation (RAG) Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN Safety & Evaluation
    Sharp Rates and a One-Line Correction for Spectral Representation Learning
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Agents & Tool Use
    Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control
    AI Agents Neural Network Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Industry Adoption
    KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI
    Deep Learning Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Look Before You Leap: Factual Decoding with Internal Attribution Signals
    Inference Machine Learning Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.LG (Machine Learning) · EN Training & Fine-tuning
    Principal-timestep Restricted Init via Sparse Matrix-decomposition in Flow-matching
    Fine-tuning Neural Network
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.LG (Machine Learning) · EN New Model Releases
    The Misery of Mechanistic Interpretability: A Formal Perspective
    GPT Llama Neural Network
    Read original (arXiv cs.LG (Machine Learning)) ↗
  • arXiv cs.CL (Computation and Language) · EN Safety & Evaluation
    Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models
    Reinforcement Learning from Human Feedback (RLHF) Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • IEEE Spectrum (AI section) · EN Agents & Tool Use
    Why Andon Labs Puts AI Agents in Charge of Real Businesses
    Andon Labs puts AI agents in charge of real stores to measure autonomy
    AI Agents Anthropic Claude Google OpenAI
    AI safety firm Andon Labs runs a real San Francisco store and cafe managed by AI agents to measure how much real-world responsibility they can handle. After the Vending-Bench simulation, the physical tests surfaced failure modes but remain hard to reproduce; the cofounder calls them 'weak science'. Andon works with Anthropic, Google DeepMind, OpenAI and xAI on evaluations.
    Read original (IEEE Spectrum (AI section)) ↗
  • Cohere Blog · EN Safety & Evaluation
    Who Gets to Define the Rules for AI?
    Cohere CEO calls Anthropic's antitrust-waiver AI safety plan a 'cartel'
    Reinforcement Learning
    Aidan Gomez argues Dario Amodei's antitrust-waiver plan letting top labs set AI safety rules would entrench incumbents, as rating agencies once did. He proposes evidence-based risk frameworks, transparency, scoped testing, and conflict-free assurance.
    Read original (Cohere Blog) ↗
  • arXiv cs.CL (Computation and Language) · EN Safety & Evaluation
    SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
    Retrieval-Augmented Generation (RAG) Transformer
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • Simon Willison's Weblog · EN Safety & Evaluation
    Quoting Boris Cherny
    Boris Cherny: Claude-written production code needs a higher bar
    AI Agents Anthropic Claude Neural Network
    Simon Willison quotes Anthropic's Boris Cherny arguing that production code written by Claude should be held to a higher bar than human-written code. Anthropic layers on guardrails: extensive lint rules and tests, Claude-driven end-to-end tests, Claude-powered fuzzers run daily, automated code and security reviews, and automated refactoring. Without them, he warns, the result is hard to maintain.
    Read original (Simon Willison's Weblog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
    Retrieval-Augmented Generation (RAG) Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Safety & Evaluation
    ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC
    Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    MAxBench: A Multinomial Concept Recovery Benchmark
    Meta
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Safety & Evaluation
    Unified CT and MRI Pancreas Segmentation for Label-Efficient Cross-Modality Subregion Transfer
    Neural Network Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Safety & Evaluation
    Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems
    Robotics
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Safety & Evaluation
    3D CT-to-PET Translation via Latent Brownian Bridge Diffusion
    Deep Learning Meta Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK
    Inference Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    Online Video Agent Harness for Long Video Understanding
    AI Agents Deep Learning Neural Network Software Engineering Speech Processing
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    Assisted Spatial Cognition Through Vision-Language Models
    Computer Vision Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Safety & Evaluation
    When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation
    Health & Bio
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗