codefuse-ai/awesome-code-llm

[TMLR] A curated list of language modeling researches for code (and other software engineering activities), plus related datasets.

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 18 minutes ago
Added to GitGenius on September 20th, 2026
Created on September 21st, 2023
Open Issues & Pull Requests: 21 (+0)
GitHub issues: Enabled
Number of forks: 239
Total Stargazers: 3,442 (+0)
Total Subscribers: 90 (+0)

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 12
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 197 days
Stale 30+ days: 11
Stale 90+ days: 8

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • OSS-Easy (4)

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Awesome-Code-LLM is a curated research collection that catalogs language modeling studies focused on code and software engineering tasks, along with associated datasets.

The collection addresses the need for a comprehensive overview of the rapidly expanding field of code-focused language models. It organizes research across multiple dimensions: foundational models and pretraining approaches, adaptation of existing language models to code, instruction fine-tuning and reinforcement learning techniques, the intersection of coding with reasoning tasks, and specialized applications for low-resource or domain-specific languages. The repository structures papers into categories including code search, code generation, code understanding, and code agents, making it possible to navigate the landscape of code LLM research systematically.

Developers and researchers building code-related AI systems should use this collection to understand the state of the field and identify relevant prior work. It suits anyone working on code generation, code search, code understanding, or reasoning systems that leverage code. The repository is particularly valuable for those exploring how existing language models can be adapted for code tasks, or for understanding specialized applications like code agents and interactive coding systems. The collection includes papers from major venues and conferences, providing access to both established and emerging research directions.

The project maintains active curation with regular additions of papers from major conferences and venues. Contributors are invited to report missing papers, miscategorizations, or incomplete reference information through the issue tracker. The repository includes supplementary resources beyond papers, such as links to code implementations and model releases associated with featured research.