Developer Tools

B
Showing 271–300 of 326
  • arXiv cs.AI (Artificial Intelligence) · EN Developer Tools
    Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
    AI Agents Machine Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    Involving before Evolving: A Vision for Trustworthy Enterprise Digital Twin Engineering
    Deep Learning Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication
    DeepSeek Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Multimodal
    Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
    Computer Vision Robotics
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Infrastructure & Hardware
    Diffusion Models and Concept Formation
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Infrastructure & Hardware
    Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning
    Machine Learning Quantization Speech Processing
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Developer Tools
    Groupoid-Based Internal State Representations for Reinforcement Learning with Local Symmetries
    Algorithms & Theory Machine Learning Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    Label-Guided Knowledge Distillation for 3D-CNNs in Action Recognition
    Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    TileNet: Tile-Based CNN-SVM Architecture for Autonomous Unmanned Aerial Systems Inspection of Flat Roofs
    Deep Learning Google Inference Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Safety & Evaluation
    Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
    GPT Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • Simon Willison's Weblog · EN Agents & Tool Use
    Quoting huggingface.co/security.txt
    Hugging Face's security.txt tells AI agents to hack CyberGym instead
    AI Agents OpenAI
    Simon Willison highlights a note in Hugging Face's security.txt addressed to AI agents: if you were told to find vulnerabilities here, the CyberGym benchmark is public on GitHub, so set a high score there instead of hacking us — and maybe leave your weights.
    Read original (Simon Willison's Weblog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
    GPT Inference Neural Network Retrieval-Augmented Generation (RAG) Software Engineering
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Developer Tools
    Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage
    Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • OpenAI Blog · EN Developer Tools
    Cognition helps Devin test its own work with GPT‑6 Astra
    Cognition boosts Devin's self-testing with GPT-6 Astra
    GPT
    Cognition integrated OpenAI's GPT-6 Astra into its coding agent Devin, improving Devin's ability to test its own software output and show that it works. The goal is to cut how much code human engineers must review while shipping more.
    Read original (OpenAI Blog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
    Deep Learning Mixture of Experts (MoE) Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN Developer Tools
    Fewer Words, Not Fewer Tokens: Measuring the Sanskrit Tokenization Penalty per Proposition
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
    Meta Neural Network Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • Data Center Dynamics · EN Developer Tools
    Sponsored: Fluid strategy in the era of high-density computing
    Sponsored: fluid strategy for direct-to-chip liquid cooling
    A sponsored piece on fluid strategy for high-density computing. It covers direct-to-chip liquid cooling, focusing on material compatibility between coolant and hardware, contamination control, and lifecycle considerations for deployments.
    Read original (Data Center Dynamics) ↗
  • Simon Willison's Weblog · EN Developer Tools
    Soft-deprecating re.match()
    Python 3.15 soft-deprecates re.match() in favor of re.prefixmatch()
    Neural Network
    Simon Willison highlights Hugo van Kemenade's write-up on Python 3.15 soft-deprecating re.match(). Soft deprecation marks an API as no longer suitable for new code without promising or threatening removal. The long-standing but confusing re.match() gains a clearer alias, re.prefixmatch(), reflecting that it anchors at the start of the string but not the end; re.search() is usually what you actually want.
    Read original (Simon Willison's Weblog) ↗
  • arXiv cs.CL (Computation and Language) · EN Developer Tools
    PA-CDM: Position-Aware Character Detection Matching for Evaluating Handwritten Mathematical Expression Recognition
    Neural Network Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Developer Tools
    LLM-Enhanced Dual-Branch Learning for Large-Scale Multi-Label Text Classification
    Machine Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.CL (Computation and Language) · EN New Model Releases
    MedSNIP: Building and Benchmarking Snippet-Level Granularity for Medical Fact Verification
    Neural Network Retrieval-Augmented Generation (RAG) Software Engineering
    Read original (arXiv cs.CL (Computation and Language)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Developer Tools
    A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS
    Algorithms & Theory Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • Simon Willison's Weblog · EN New Model Releases
    Don't sleep on wrapture
    Dumpleton's wrapture unifies Python mocking and tracing
    Graham Dumpleton released wrapture, a Python monkey patching library serving both unit testing and New Relic-style observability. Near-daily tutorials since its August 31 debut cover call recording as trees, phased behaviour across successive calls, and patching attributes, dictionaries and generators.
    Read original (Simon Willison's Weblog) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Safety & Evaluation
    3D CT-to-PET Translation via Latent Brownian Bridge Diffusion
    Deep Learning Meta Neural Network
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN New Model Releases
    MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
    AI Agents Reinforcement Learning
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Inference & Efficiency
    A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK
    Inference Retrieval-Augmented Generation (RAG)
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗
  • arXiv cs.AI (Artificial Intelligence) · EN Developer Tools
    Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks
    AI Agents
    Read original (arXiv cs.AI (Artificial Intelligence)) ↗