• Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication
    Sep 16 2026
    Hospital discharge summaries often blur when things happened. This framework tags every event in the note with an ID and ties it back to timestamped rows in the patient's records. On 40 critical-care summaries, it recovered 43% more events and came close to clinician annotations. The team also built GAVEL, a model that judges competing timelines, and human reviewers upheld most of its calls. Authors: Sayantan Kumar, Nicolas Grimaldi, Jack Cummins, Jeremy C. Weiss Paper: https://arxiv.org/abs/2609.13062v1
    Mostra di più Mostra meno
    2 min
  • Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction
    Sep 16 2026
    Building benchmarks for robots and other embodied AI is slow, and an early mistake can quietly spoil the result. Embodied-BenchForge has agents build the benchmark, check each intermediate piece, and redo only the parts that fail. It produced six question-answering benchmarks and one interactive benchmark with 220 tasks. Authors: Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Zhe Ji, Jinshan Lai, Xi Ren, Danyang Li, Zheng Yang, Jianwei Hu, Qiang Ma Paper: https://arxiv.org/abs/2609.13082v1
    Mostra di più Mostra meno
    3 min
  • Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
    Sep 16 2026
    One diffusion model treats language, camera views, goals and robot actions as the same kind of token. That lets it predict actions, the next view and the end state with a single model. Pretrained on about 1.33 million robot trajectories, it averaged 78.4% success on a real Franka arm across four conditions. A faster implementation cut action decoding time by up to 29 times. Authors: Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do Paper: https://arxiv.org/abs/2609.13053v1
    Mostra di più Mostra meno
    3 min
  • Attention Quantization for Tabular Foundation Models
    Sep 16 2026
    Tabular foundation models are served differently from chatbots, so the usual speedups don't carry over. Here the bottleneck is the attention calculation. Converting queries, keys and values to FP8 gave up to 1.7 times the speed with no meaningful accuracy loss on TabPFN-v3 and TabICLv2. The catch: quantization error on test rows has to match the training rows, or accuracy drops sharply. Authors: Jonas M. Kübler, Benjamin Jäger, Klemens Flöge, Noah Hollmann, Frank Hutter Paper: https://arxiv.org/abs/2609.13031v1
    Mostra di più Mostra meno
    3 min
  • Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies
    Sep 16 2026
    Driving policies trained in simulators can inflate their safety scores with sudden, jerky moves no passenger would accept. This paper redraws the set of allowed controls at every step so they stay inside comfort limits, which tighten as speed rises. On Waymo driving data, comfort violations stayed under 1% and the policy still navigated better than the baselines. Authors: Anna Rothenhäusler, Daniel Jost, Raghu Rajan, Faris Janjos, Oliver Scheel, Andreas Look, Joschka Boedecker Paper: https://arxiv.org/abs/2609.13011v1
    Mostra di più Mostra meno
    3 min
  • MAxBench: A Multinomial Concept Recovery Benchmark
    Sep 16 2026
    Research on steering language models mostly deals with yes-or-no concepts such as refusal. MAxBench tests concepts with many categories, such as animals or countries. Across 10 methods, 6 concepts and 4 models, affine subspaces worked best, largely because of better offsets. No method beat plain prompting. Authors: Divya Appapogu, Freya Behrens, Yonatan Belinkov, Aaron Mueller Paper: https://arxiv.org/abs/2609.13072v1
    Mostra di più Mostra meno
    2 min
  • MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
    Sep 16 2026
    Voice agents are mostly tested in one-on-one conversation. MP-Bench tests them in groups. Across 12 agents, real-time systems scored 22% or less on comprehension and did no better than chance at knowing when to speak. Authors: Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee, David Harwath Paper: https://arxiv.org/abs/2609.13076v1
    Mostra di più Mostra meno
    2 min
  • Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
    Sep 16 2026
    Can AI research agents handle a messy, open-ended industry problem? This study put them on telecom ticket retrieval. With little supervision they reached 90% of the best human result in 10 weeks instead of 10 months, at up to $200 a run. They were good at tuning but short on intuition, and the authors recommend pairing them with human researchers. Authors: Junghyun Min, Huseyin Uzunalioglu, Mohamed Trabelsi Paper: https://arxiv.org/abs/2609.13073v1
    Mostra di più Mostra meno
    2 min