← Today · Tue, Oct 6

SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

We ve been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project

At a glance

  • reddit.com: SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]
  • arxiv.org: Benchmarks in Leipzig
  • arxiv.org: MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

The story

reddit.com: We ve been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project s own tests, in a container with no network, and the repo is cut down to a single commit so the agent can t recover the fix from git history. https://preview.redd.it/0uopztlmpsth1.png?width=1200 format=png auto=webp s=59c0698ca0a76c17be41e41362324d6a262f1f25 Some findings: With one attempt per task GLM-5.3 Flash scored 85%. With two to three attempts it scored 82%, within the margin of error of GPT-5.6 Luna (81%). The leaderboard now shows the number of attempts and the interval for every score. https://preview.redd.it/9dqmk9uppsth1.png?width=1200 format=png auto=webp s=006cd587d11c03e47baf50c9954d95d937d09005 About half the tasks are easy for every model (near 100%). The other half is where they actually differ: 50%, 45% and 23% on the hard ones. Most of the difference between models comes from the hard half. Since every fix is public on GitHub, we reviewed all 11k commands the agents ran. 69 tried to access the network and all failed. GLM tried 50 times to pip download the already-fixed release of the library it was fixing. We also checked contamination by comparing older bugs (pre-2026) with newer ones of similar size. Older ones are solved about 9 points more often, but the confidence interval crosses zero, so we can t say much yet. Half the tasks are private. So far public and private scores line up for all three models. Results and every agent run: https://labs.evaligo.com/swe-race?utm_source=reddit utm_medium=ml utm_campaign=launch Tasks: https://huggingface.co/datasets/evaligo/swe-race The protocol follows DeepSWE (100 steps). Feedback on it, and suggestions for which models to run next, are welcome. submitted by /u/heyitsdannyle [link] [comments]

arxiv.org: arXiv:2606.05818v2 Announce Type: replace-cross Abstract: Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the three-day workshop Benchmarks in Leipzig with 35 participants at the Max Planck Institute for Mathematics in the Sciences in Leipzig, Germany. We present the resulting collection of 100~questions. We evaluated these questions in three stages: a single attempt by five state-of-the-art LLMs and their predecessors, followed by a 20-runs-per-model evaluation with three of these models, and finally a 3-run attempt with two heavy-thinking models. After Stage 1, 41 questions remained completely unsolved; after Stage 2, this count dropped to 16; and we concluded Stage 3 with only 2 unsolved questions. This demonstrates that the mathematical reasoning capabilities of LLMs are becoming impressive. In September 2026, we added a fourth stage in which the next generation of models attempted all 100 questions once more, after which only 1 question remains unsolved.

arxiv.org: arXiv:2606.06696v2 Announce Type: replace-cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy. Realizing this potential, however, requires robust and fine-grained visual perception. Models need to correctly interpret subtle features in images, and they must do so across diverse biomedical modalities, scales, and contexts. Nevertheless, current benchmarks remain limited. To address these gaps, we introduce the Massive Multimodal Biomedical Understanding (MMBU) benchmark. It is the largest biomedical vision and language benchmark to date, covering 35 submodalities with rich structured metadata. It includes both open and closed versions of ungrounded classification, grounded classification, and object detection, enabling systematic evaluation of model performance across biological scales, clinical settings, and imaging modalities. Evaluating 16 open-weight and 5 frontier VLMs in the main comparison, we find that while medical adaptation provides measurable gains for some models, the high accuracy often reported on established benchmarks can mask deficiencies in visual perception and domain generalization. We further define an open-ended MMBU hard set, split into a released public subset and a held-out private subset, to stress-test frontier models.

arxiv.org: arXiv:2607.00218v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) are increasingly proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming activity, a distinction obscured by binary safety benchmarks. We introduce EgoSafetyBench, an egocentric video benchmark of 1,200 robot-view scenarios annotated at half-second granularity, with two evaluation tracks. The situational track (800 scenarios) spans routine, safe-but-suspicious, obvious-hazard, and contextual-hazard scenes. The visual-channel track (400 scenarios) tests whether misleading in-scene text corrupts physical-safety judgments, using matched truthful controls. Both tracks use contrastive ladders: near-identical scenarios differing in a single visible deciding cue, forcing predictions to hinge on that cue. Across ten open- and closed-source VLMs, we find that guards often recognize videos containing hazards yet miss the specific hazardous moments, especially for contextual hazards. Misleading in-scene signs further degrade all tested guards: vulnerable models miss up to a third of hazards, while seemingly robust models often over-intervene on safe content. Matched controls show that apparent robustness can reflect indiscriminate alarming rather than true physical reasoning. A 20-clip real-video sanity check further shows that model rankings transfer beyond synthetic rendering, with Spearman \r{ho} = 0.87 for hazard miss rate and 0.94 for visual-channel mismatch recall.

Get tomorrow's scan at 7am

The same ranked list, in your inbox. Nothing else, ever.

← Back to Today