BBVA signaled it would place OpenAI at the core of its banking operations. All five sources are OpenAI official—a single-vendor framing in which a large financial-institution deployment is presented as an announcement, driven by business adoption rather than research. The focus is embedding generative AI broadly across internal processes and customer service, elevating it from pilots to core systems; the through-line is less raw model performance than a regulated bank's stance of treating AI as 'mission-critical.' It stands out as full-scale adoption in a conservative domain. But this is a statement of intent—implementation scope, risk management, and results remain to confirm.
BBVA puts OpenAI at banking core
BBVA puts OpenAI at banking core
OpenAI WebRTC Audio Session, now with document context
Simon Willison adds document context to his OpenAI WebRTC audio tool
Simon Willison shares a quote from Andrew Singleton
スマホからWindowsのCodexアプリを操作できるの? 外出中でもAIコーディングを止めない方法
OpenAI Codex on Windows now controllable from smartphones
Timing Trick Cuts Energy Used in LLM Training by Up to 14 Percent
Timing trick cuts LLM training energy use by up to 14 percent
OpenAI confidentially files for a US IPO; timing undecided
Academic (arxiv etc.) 42 ▾
HumP-KD: uncertainty-aware distillation for efficient fire classification
A longitudinal taxonomy of silent failures in a production LLM agent runtime
A temporal planning framework for disruption-aware railway routing
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
From shield to target: DoS attacks on LLM-based agent guardrails
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Every Eval Ever: a unifying schema and repository for AI evaluations
Fodor and Pylyshyn's Systematicity Challenge Still Stands
Fodor and Pylyshyn's systematicity challenge still stands
tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration
tap: a file-based protocol for heterogeneous LLM agent collaboration
Retrospective Progress-Aware Self-Refinement for LLM Agent Training
Progress-aware self-refinement for training LLM agents
Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge
Does the LLM judge prefer English? Testing language-switching invariance
CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward
CacheRL trains tool-calling agents via cached rollouts and hybrid reward
Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
Spatio-temporal audio language modeling for dynamic sound sources
HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents
HyperTool folds tool workflows into code, lifting MCP-Universe to 35.29%
Recursive Agent Harnesses lift long-context coding accuracy to 81.36%
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
AgentBeats proposes agentified, protocol-standardized agent assessment (AAA)
EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis
EpiBench: no AI agent passes a majority of epigenomics analysis tasks
A Three-Layer Framework for AI in Scientific Discovery
A three-layer view of AI in discovery centers on model formation (Layer 2)
AgentRivet: an automated system for producing Rivet routines from journal publications
AgentRivet auto-generates missing Rivet routines from physics papers
CRAFTIIF: unsupervised interpretable detection of four anomaly types
'Evaluation sovereignty' exposes label-authority bias in classification metrics
Optimizing Appliance Scheduling for Solar Energy Management Using Metaheuristic Algorithms
Metaheuristics optimize appliance scheduling to maximize solar self-consumption
Layer-Resolved Optimal Transport for Hallucination Detection in NMT and Abstractive Summarization
Layer-wise optimal transport probes hallucination detection in NMT
SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection
SICI complexity index reveals regime shifts in LLM stance detection
HyPE encodes persona relations as hypergraphs for grounded dialogue
TAHOE: Text-to-SQL with Automated Hint Optimization from Experience
Tahoe learns hints from experience to optimize production Text-to-SQL
Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy
Atlas H&E-TME profiles tissue from H&E slides at pathologist-level accuracy
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Claw-SWE-Bench fairly benchmarks OpenClaw-style coding agent harnesses
Adjoint Method versus Physics-Informed Neural Networks in PDE-Constrained Inverse Problems
Adjoint optimization vs PINNs: a fair test on PDE inverse problems
SpikeDecoder: Realizing the GPT Architecture with Spiking Neural Networks
SpikeDecoder realizes a GPT-style decoder with spiking neural networks
Finding Multiple Interpretations in Datasets
Method finds equally accurate models with divergent interpretations
Market Design for AI: Beyond the Copyright Binary
Market design for AI training data beyond the copyright binary
Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation
Soft-prompt tuning enables fair, efficient LLM benchmark evaluation
Debiasing Without Protected Attributes: Latent Concept Erasure from Textual Profiles
H-SAL debiases LLMs without access to protected attributes
M-EDESConv corpus boosts multilingual emotional validation in dialogue
ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
ABC-Bench evaluates LLM agents' biosecurity-relevant capabilities
Generative Archetype-Grounded Item Representations for Sequential Recommendation
GenAIR: archetype-grounded item representations for sequential recommendation
Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages
Frontier coding agents use metaprogramming for esoteric languages
It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO
One-shot GRPO on a single biased example breaks LLM guardrails
Ethical and Technical Limits of Deepfake Speech Datasets
Auditing 39 deepfake speech datasets reveals fairness gaps and overlap
Human-AI Teaming Through the Lens of Calibration
A calibration lens on human-AI teaming: combination breaks calibration
Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory
Infini Memory: topic documents for long-term LLM agent memory
Speaker Group Encoding in Self-supervised Speech Recognition Models
How self-supervised speech models encode speaker group traits
Causal Ensemble Agent: Hierarchical Causal Discovery with LLM-guided Expert Reweighting
LLM referee reweights expert ensemble for causal discovery