OpenAI unveiled an autonomous 'AI chemist' aimed at improving drug synthesis. The signal is announcement-led—an official post followed by ITmedia coverage and community reaction on lobste.rs and HN, ahead of any academic papers. The direction is automating lab work, handing route search and optimization to an agent, and the notable part is bringing agents into drug discovery as an applied domain. But this is the announcement stage: reproducibility in real experimental settings, success rates, and integration into existing workflows are unconfirmed—grounded validation, not flash, is what to watch next.
OpenAI debuts autonomous AI chemist
OpenAI debuts autonomous AI chemist
A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry
OpenAI's near-autonomous AI chemist improves a key medicinal reaction
「AIを使う学生」vs.「使わない学生」、エッセイが創造的なのはどっち? 米大学が2025年に実証実験
AI-using vs non-using students: whose essays are more creative?
GPT‑NL: a sovereign language model for the Netherlands
GPT-NL: a sovereign language model for the Netherlands
OpenAIの高度AIでソフトバンクの脆弱性を1万件発見 孫正義氏「大変な危機」 日本の重要インフラ企業へ診断サービス提供
SoftBank unveils OpenAI-powered Patching-as-a-Service security offering
CrankGPT — Local Human-powered AI
CrankGPT pitches a tongue-in-cheek local, human-powered 'AI'
Academic (arxiv etc.) 20 ▾
Explaining Attention with Program Synthesis
Explaining attention via program synthesis for interpretability
A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2
A multi-domain benchmark to detect GPT-Image-2 text-rich images
TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology
TxBench-PP evaluates AI agents on preclinical pharmacology
CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System
CAPRA: a multi-agent LLM system for software architecture feedback
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
ReproRepo scales reproducibility audits using GitHub repo issues
RubricsTree: scalable open-ended evaluation of personal health agents
An agentic benchmark for implicit animal welfare in frontier AI
Structural role injection in Handlebars-templated LLM prompts
Querying an astronomical database using large language models: the ALeRCE text-to-SQL system
A text-to-SQL system for querying the ALeRCE astronomical database
Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
Security and privacy prompts in the wild: what users ask LLMs
Synthetic lived experience in AI peer-like caregiver support
Toward Accessible Psychotherapy Training Using AI-Driven Interactive Patient Avatars
AI-driven patient avatars for more accessible psychotherapy training
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
From trainee to trainer: LLM-designed RL training environments
Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models
Binary Tracking: open vision-language models for spatial QA and navigation
Reasoning hop-count predicts clinical AI failure in EHR QA
Paper: framework measures LLM search-agent endorsement risk
LLM-based Visual Code Completion for Aerospace Geometric Design
Paper: LLM visual-programming copilot for aerospace design
Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents
Paper on evaluator preference collapse in self-evolving agents
Paper frames LLM sycophancy as material failure (title only)
BD-LSC: a new benchmark dataset for lexical semantic change detection