CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding Capabilities of CodeLLMs Paper • 2410.01999 • Published Oct 2, 2024 • 10
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale Paper • 2409.16299 • Published Sep 9, 2024 • 11
XMainframe: A Large Language Model for Mainframe Modernization Paper • 2408.04660 • Published Aug 5, 2024
Building AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learned Paper • 2603.05344 • Published Mar 5 • 7
TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation Paper • 2606.13714 • Published Jul 27
COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation Paper • 2604.03986 • Published Apr 5
CodeWiki: Evaluating AI's Ability to Generate Holistic Documentation for Large-Scale Codebases Paper • 2510.24428 • Published Oct 28, 2025 • 4
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios Paper • 2512.18470 • Published Dec 20, 2025 • 12
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating Paper • 2505.10860 • Published May 16, 2025 • 1
SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs Paper • 2504.14757 • Published Apr 20, 2025
Building AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learned Paper • 2603.05344 • Published Mar 5 • 7
Building AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learned Paper • 2603.05344 • Published Mar 5 • 7
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios Paper • 2512.18470 • Published Dec 20, 2025 • 12
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios Paper • 2512.18470 • Published Dec 20, 2025 • 12
CodeWiki: Evaluating AI's Ability to Generate Holistic Documentation for Large-Scale Codebases Paper • 2510.24428 • Published Oct 28, 2025 • 4