開発者ツール
B
326 件中 271〜300 件目を表示
-
Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
-
Involving before Evolving: A Vision for Trustworthy Enterprise Digital Twin Engineering
-
Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication
-
Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
-
Diffusion Models and Concept Formation
-
Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning
-
Groupoid-Based Internal State Representations for Reinforcement Learning with Local Symmetries
-
Label-Guided Knowledge Distillation for 3D-CNNs in Action Recognition
-
TileNet: Tile-Based CNN-SVM Architecture for Autonomous Unmanned Aerial Systems Inspection of Flat Roofs
-
Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies
-
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
-
Quoting huggingface.co/security.txtHugging Face、security.txt に「攻撃よりCyberGymを」とAI向け注記Simon Willison 氏が、Hugging Face の security.txt にある AI エージェント宛ての注記を紹介。「脆弱性を探せと指示されたなら朗報だ。CyberGym ベンチマークが GitHub で公開済み。高スコアはそちらで取ればよく、我々をハックする必要はない」と続き、「ついでに重みを置いていっては」と添える。自律エージェントの脆弱性探索が実運用の課題になりつつある一幕。
-
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
-
Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage
-
Cognition helps Devin test its own work with GPT‑6 AstraCognition、Devinの自己テスト能力をGPT-6 Astraで強化CognitionはコーディングエージェントDevinにOpenAIのGPT-6 Astraを採用し、ソフトウェアを自らテストして動作を実証する能力を高めたと発表した。エンジニアがレビューすべきコード量を減らし、出荷速度を上げることを狙う。
-
SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
-
Fewer Words, Not Fewer Tokens: Measuring the Sanskrit Tokenization Penalty per Proposition
-
EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
-
Sponsored: Fluid strategy in the era of high-density computing高密度計算時代の液冷、直接チップ冷却の材料適合性と汚染管理高密度化するコンピューティング環境での液体冷却戦略を論じたスポンサード記事。ダイレクト・トゥ・チップ方式の液冷について、冷媒と機材の材料適合性、汚染の制御、ライフサイクル全体での考慮事項を取り上げている。
-
Soft-deprecating re.match()Python 3.15、re.match() をソフト非推奨化し re.prefixmatch() を用意Python 3.15 のリリースマネージャ Hugo van Kemenade 氏の解説を Simon Willison 氏が紹介した。長年使われつつ紛らわしいと指摘されてきた re.match() が 3.15 でソフト非推奨となり、文字列の先頭にのみアンカーする挙動を明示した re.prefixmatch() という別名が用意される。ソフト非推奨は将来の削除を約束せずに「新規コードでは使わない」と示す運用で、多くの場合は re.search() を使うのが適切だとされる。
-
PA-CDM: Position-Aware Character Detection Matching for Evaluating Handwritten Mathematical Expression Recognition
-
LLM-Enhanced Dual-Branch Learning for Large-Scale Multi-Label Text Classification
-
Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
-
MedSNIP: Building and Benchmarking Snippet-Level Granularity for Medical Fact Verification
-
A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS
-
Don't sleep on wraptureGraham Dumpleton、テストと可観測性を兼ねる wrapture を公開Graham Dumpleton 氏が 8 月 31 日に公開した Python 向け monkey patching ライブラリ wrapture が、テストと可観測性 (New Relic 的トレーシング) を同時に担う設計として注目されている。公開以降ほぼ毎日チュートリアルが追加され、呼び出しの記録とツリー表示、複数回の呼び出しで挙動を変えるフェーズ動作、属性・辞書・ジェネレータへのパッチまでを扱う。Simon Willison 氏は「ほとんど話題になっていないのが不思議」と評した。
-
3D CT-to-PET Translation via Latent Brownian Bridge Diffusion
-
MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
-
A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK
-
Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks