Cohere published techniques to improve fairness in LLM serving (inference). The composition is strongly academic—one official Cohere source and four arXiv papers—so this reads less as a product launch than a research-backed argument about inference quality. It addresses correcting latency and resource-allocation skew when handling concurrent requests to keep service fair across users—an operational but quietly fundamental theme. The through-line is a shift of attention from scaling models up to serving the same model fairly and efficiently. It's still research-led; real-service impact and adoption are yet to confirm.
Cohere tackles LLM serving fairness
Cohere tackles LLM serving fairness
LLM Serving Fairness: No more noisy neighbours
Cohere ensures fair compute sharing across LLM serving tenants
Academic (arxiv etc.) 21 ▾
Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
LOCUS releases a US local-ordinance corpus for legal AI
UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning
UBP2: uncertainty-balanced planning for efficient preference-based RL
IndicContextEval: audio-LLM context use across 8 Indic languages
A Technical Taxonomy of LLM Agent Communication Protocols
A technical taxonomy of LLM agent communication protocols
Towards an Agent-First Web: Redesigning the Web for AI Agents
Towards an agent-first web: redesigning the web for AI agents
Which paper sections best reveal research methods?
Improving Medical Communication using Rubric-Guided Counterfactual Recommendations
Rubric-guided counterfactual recommendations for medical communication
RubricsTree: scalable open-ended evaluation of personal health agents
Descriptor: Certus Caliber Classification Gunshot Dataset (C3GD)
C3GD: a public field-collected gunshot muzzle-blast sound dataset
Knowledge Reutilization in Meta-Reinforcement Learning
A meta-knowledge reutilization framework for meta-RL across agents
Unintended Effects of Geographic Conditioning in Large Language Models
Unintended regional biases from geographic conditioning in LLMs
Learning Fair Pareto-Optimal Policies in Multi-Objective Reinforcement Learning
Learning fair Pareto-optimal policies in multi-objective RL
Meta-classification of one-class models via ranking correlation and kNN
WallZero: Mastering the Game of WallGo with Strategic Analysis
WallZero masters the board game WallGo with strategic analysis
When Multiple Scripts Matter: Evaluating ASR in Clinical Settings
Evaluating ASR in clinical settings when multiple scripts matter
Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns
Reusing web skills via transferable interaction patterns
Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio
A benchmark for LLM agents on Nature Portfolio meta-analyses
Phase, not magnitude, carries identity inside image classifiers
RAID: Semantic Graph Diffusion for True Cold-Start and Cross-Lingual Forecasting
RAID: retrieval-augmented diffusion for cold-start, cross-lingual forecasting
Benchmark suite for federated noisy-label medical image segmentation
FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection
FraudSMSWalker benchmark targets URL-masked SMS-to-webpage fraud