OpenAI × New Model Releases

OpenAI debuts autonomous AI chemist

OpenAI debuts autonomous AI chemist

✎ Story body

OpenAI unveiled an autonomous 'AI chemist' aimed at improving drug synthesis. The signal is announcement-led—an official post followed by ITmedia coverage and community reaction on lobste.rs and HN, ahead of any academic papers. The direction is automating lab work, handing route search and optimization to an agent, and the notable part is bringing agents into drug discovery as an applied domain. But this is the announcement stage: reproducibility in real experimental settings, success rates, and integration into existing workflows are unconfirmed—grounded validation, not flash, is what to watch next.

▲ Official & Press
Official

A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry

OpenAI Blog ・ 2026-06-17 ・ 📌

OpenAI's near-autonomous AI chemist improves a key medicinal reaction

Press

「AIを使う学生」vs.「使わない学生」、エッセイが創造的なのはどっち? 米大学が2025年に実証実験

ITmedia AI+ ・ 2026-06-17

AI-using vs non-using students: whose essays are more creative?

Community

GPT‑NL: a sovereign language model for the Netherlands

Hacker News (Front Page) ・ 2026-06-16

GPT-NL: a sovereign language model for the Netherlands

Press

OpenAIの高度AIでソフトバンクの脆弱性を1万件発見 孫正義氏「大変な危機」 日本の重要インフラ企業へ診断サービス提供

ITmedia AI+ ・ 2026-06-16

SoftBank unveils OpenAI-powered Patching-as-a-Service security offering

Community

CrankGPT — Local Human-powered AI

Lobste.rs (AI tagged) ・ 2026-06-15

CrankGPT pitches a tongue-in-cheek local, human-powered 'AI'

Academic (arxiv etc.) 20 ▾
Academic

Explaining Attention with Program Synthesis

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

Explaining attention via program synthesis for interpretability

Academic

A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

A multi-domain benchmark to detect GPT-Image-2 text-rich images

Academic

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

TxBench-PP evaluates AI agents on preclinical pharmacology

Academic

CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-17

CAPRA: a multi-agent LLM system for software architecture feedback

Academic

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues

arXiv cs.CL (Computation and Language) ・ 2026-06-16

ReproRepo scales reproducibility audits using GitHub repo issues

Academic

RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

arXiv cs.CL (Computation and Language) ・ 2026-06-16

RubricsTree: scalable open-ended evaluation of personal health agents

Academic

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

arXiv cs.CL (Computation and Language) ・ 2026-06-16

An agentic benchmark for implicit animal welfare in frontier AI

Academic

Structural Role Injection in Handlebars-Templated LLM Prompts: Triple-Brace Interpolation, Delimiter Family, and the Limits of HTML Auto-Escaping

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Structural role injection in Handlebars-templated LLM prompts

Academic

Querying an astronomical database using large language models: the ALeRCE text-to-SQL system

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-16

A text-to-SQL system for querying the ALeRCE astronomical database

Academic

Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Security and privacy prompts in the wild: what users ask LLMs

Academic

When AI Says "I have been in similar situations": Synthetic Lived Experience in Peer-Like Caregiver Support

arXiv cs.CL (Computation and Language) ・ 2026-06-16

Synthetic lived experience in AI peer-like caregiver support

Academic

Toward Accessible Psychotherapy Training Using AI-Driven Interactive Patient Avatars

arXiv cs.CL (Computation and Language) ・ 2026-06-16

AI-driven patient avatars for more accessible psychotherapy training

Academic

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning

arXiv cs.CL (Computation and Language) ・ 2026-06-16

From trainee to trainer: LLM-designed RL training environments

Academic

Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models

arXiv cs.AI (Artificial Intelligence) ・ 2026-06-15

Binary Tracking: open vision-language models for spatial QA and navigation

Academic

Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Reasoning hop-count predicts clinical AI failure in EHR QA

Academic

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Paper: framework measures LLM search-agent endorsement risk

Academic

LLM-based Visual Code Completion for Aerospace Geometric Design

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Paper: LLM visual-programming copilot for aerospace design

Academic

Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Paper on evaluator preference collapse in self-evolving agents

Academic

Sycophancy as Material Failure under Pushback Loading: A Multi-Axis Characterization Across Three Loading Cases and up to Seventeen Material Charges

arXiv cs.CL (Computation and Language) ・ 2026-06-15

Paper frames LLM sycophancy as material failure (title only)

Academic

The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

arXiv cs.CL (Computation and Language) ・ 2026-06-15

BD-LSC: a new benchmark dataset for lexical semantic change detection

← Story Archive