Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
7811 commits
Select commit Hold shift + click to select a range
3330d12
Automated community request assignment (#5147)
Phlip79 Jun 24, 2026
311416f
Clean up training.py module header (dedupe + reorganize imports/globa…
ilml Jun 24, 2026
9038381
Thread process groups through training checkpoint paths (#5486)
yashaswikarnati Jun 24, 2026
5863721
Narrow oncall responsibilities (#5490)
Phlip79 Jun 25, 2026
1c1d6b5
Add MIMO forward step and per-token loss for hetero training (#5376)
yashaswikarnati Jun 25, 2026
ea967a7
Add Nemotron6-MoE VLM model provider for MIMO example (#5374)
yashaswikarnati Jun 25, 2026
7168714
ci: auto-retry test-data download in container-build job (#5498)
ko3n1g Jun 25, 2026
3bfd87b
Force RL inference to CP=1 (#5423)
tdene Jun 25, 2026
2a43e0d
Merge cu_seqlens across micro-batch for THD attention (#5454)
deepakn94 Jun 25, 2026
da482cf
[split 4/4] Enable DSA CP and THD hooks (#5246)
HollowMan6 Jun 25, 2026
e1b8454
Fix fused MLA down projection with tensor parallelism (#5383)
sraman-rgb Jun 25, 2026
8bafe7c
Fix NameError in is_flashinfer_min_version when check_equality=False …
adityasingh2400 Jun 26, 2026
da42015
Add hybrid FSDP unit module support (#4329)
Phlip79 Jun 26, 2026
476228d
fix: set DATA_PATH for moe-dynamic-inference recipe (#5506)
ko3n1g Jun 26, 2026
c0d7848
Add --qad-train-target {base|mtp|both} for QAD / MTP QAT (frozen-base…
yeyu-nvidia Jun 26, 2026
0552f29
[Main] Generalized fix for mxfp8 param gather (#5236)
zhongbozhu Jun 26, 2026
5949478
Add CUDA graph training iteration test (#5417)
wujingyue Jun 26, 2026
847de23
test: restore G/G + lag=19 for gpt_grpo_tp4_pp1_dp2_8b throughput tes…
lauradang Jun 26, 2026
990ced9
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jun 27, 2026
cbaa6eb
ci: cache-from a single coherent buildcache donor (#5509)
ko3n1g Jun 26, 2026
ed59a0a
ci: Use GB300 for Github CI tests (#5520)
chtruong814 Jun 27, 2026
0ff7226
ci: pin HF_HUB_CACHE to bind-mounted cache for gpt-oss-20b inference …
ko3n1g Jun 27, 2026
f88b85f
Add inter-document attention masking to GPTDataset (#5298)
deepakn94 Jun 28, 2026
25f6a09
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jun 29, 2026
522a9dd
[CI] Fix `gpt_dynamic_inference_tp2_pp2_ep2_gptoss_20b_swa` tests (#5…
asolergi-nv Jun 29, 2026
3c327f3
Implement async scheduling for dynamic inference (#5453)
lmcafee-nvidia Jun 29, 2026
b9f7fb6
build: bump transformer-engine to release_v2.16.post (#5517)
ko3n1g Jun 29, 2026
819912c
Fix `isort` target Python version (#5567)
janEbert Jun 30, 2026
7b6eb02
Fix PR template typo (#5566)
janEbert Jun 30, 2026
872442a
Deduplicate tensor-splitting utility (#5545)
anlthms Jun 30, 2026
99b56a7
Add CI duties to oncall (#5510)
Phlip79 Jun 30, 2026
817c1d5
Thread dp_cp/expt_dp process groups through checkpoint load path (#5579)
yashaswikarnati Jun 30, 2026
f285ea5
Fix TEGroupedMLP pre-backward unshard in fine-grained FSDP hooks for …
rapatel Jun 30, 2026
223e244
[training migration] Finish ModelBuilder integration (#5516)
maanug-nv Jun 30, 2026
4849143
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jul 1, 2026
a3d761b
Use NVIDIA inference credentials for Claude actions (#5589)
Phlip79 Jul 1, 2026
36d11ce
Update PR instructions (#5592)
Phlip79 Jul 1, 2026
f972da7
Update mcore skill owners (#5586)
Phlip79 Jul 1, 2026
8b3d8a5
chore: rotate oncall schedule
github-actions[bot] Jul 1, 2026
3c08255
Add /claude fix workflow for on-demand PR fixes (#4862)
Phlip79 Jul 1, 2026
4c4a8ee
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jul 2, 2026
adfe9e1
[Megatron-FSDP] MaxPoolAllocator for double-buffering hybrid architec…
cspades Jul 2, 2026
60f338a
fix(tensor_parallel): _reduce returns unreduced tensor for non-contig…
Pearblossom-M Jul 2, 2026
2b551c6
Add Auto Quantize in ModelOpt quantize example (#4821)
jenchen13 Jul 2, 2026
69c4868
[Main][feat] Support CUDA Graph capture offloading modules (#3697)
lhb8125 Jul 2, 2026
0522099
E2E heterogenous non colocated MiMo training (#5602)
yashaswikarnati Jul 2, 2026
5e4fe9b
Optimize memory usage of partial CUDA graphs (#5451)
jiemingz Jul 2, 2026
ee623a9
Document stacked dependent PR handling in split PR skill (#5496)
wujingyue Jul 2, 2026
060371c
Fix Claude reaction permissions (#5613)
Phlip79 Jul 2, 2026
6608a05
Fix smoke BERT/T5 test failures (#5629)
balasaajay Jul 2, 2026
25f6117
Add NCCL symmetric-memory staging to experimental FSDP (#5440)
wujingyue Jul 2, 2026
d89aae5
Add smoke test notification functionality and update notify script (#…
balasaajay Jul 2, 2026
4828d65
Update golden value files for GPT-3 weekly (#5459)
balasaajay Jul 2, 2026
06b07a1
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jul 4, 2026
0823c73
Ignore contributor DCO failures in Claude fix (#5625)
Phlip79 Jul 5, 2026
6856424
chore(beep boop 🤖): Bump (main) (2026-07-06)
github-actions[bot] Jul 6, 2026
959f3f1
Pre-size the all-gather buffer for inference to max capacity (#5546)
santhnm2 Jul 6, 2026
7ee524e
Scatter embeddings for sequence parallelism in standalone LM forwards…
kevalmorabia97 Jul 6, 2026
bf32f44
Fix inter-document masking crash and NaNs with TP > 1 and micro_batch…
deepakn94 Jul 6, 2026
1bb5ff5
add safe version of numpy.load (#5500)
dimapihtar Jul 7, 2026
115ce7f
Fuse shared expert MLP with grouped GEMM (#5604)
sraman-rgb Jul 7, 2026
c1560d9
MoE routing analysis and metrics capture (#5220)
mathemakitten Jul 7, 2026
c797a5e
Add cspades to oncall rotation (#5695)
Phlip79 Jul 7, 2026
5dbea46
Add microbatch context helper (#5652)
wujingyue Jul 7, 2026
68f4c62
ci: Update test configurations to unify legacy scope names (#5316)
balasaajay Jul 7, 2026
6ee356a
Fix Torch FSDP2 crash: add force_all_reduce kwarg to base finish_grad…
factnn Jul 8, 2026
72a78d6
Separate mFSDP v2 unit tests (#5640)
wujingyue Jul 8, 2026
465264b
chore: rotate oncall schedule
github-actions[bot] Jul 8, 2026
3ab71ee
test(determinism): add determinism tests (#5041)
ZhiyuLi-Nvidia Jul 8, 2026
40b1fd3
deprecate common strategy (#5160)
dimapihtar Jul 8, 2026
509efe6
ci: revert unify legacy scope names (#5316) (#5709)
ko3n1g Jul 8, 2026
65d6c23
Normalize CRLF in Claude fix commands (#5712)
Phlip79 Jul 8, 2026
2a21b8e
Remove some barriers in save_checkpoint_and_time (#5557)
shurkat-nvidia Jul 8, 2026
91fcdfe
Add FSDP NVTX annotations (#5704)
wujingyue Jul 8, 2026
2d7060f
NCCL EP support (#5129)
YangFei1990 Jul 8, 2026
64bee49
Update base image to nvcr.io/nvidia/pytorch:26.06-py3 (#5632)
balasaajay Jul 8, 2026
027fa4a
Refactor RL rollout pipeline (#5491)
lauradang Jul 8, 2026
d1cab16
[2/2] Wiring cuDNN fused DSA kernels support with THD, CP and IndexSh…
HollowMan6 Jul 8, 2026
f29c747
Triton kernels - avoid recompilation and autotuning in prod (#5608)
sidsingh-nvidia Jul 8, 2026
cf2f07d
Add NeMo Transformer audio encoder model (#5565)
yqwangustc Jul 9, 2026
a2496aa
remove deprecated modules from core/dist_checkpointing (#5134)
dimapihtar Jul 9, 2026
328a77f
fix(fsdp): import os in safe_get_rank fallback (#4959)
fallintoplace Jul 9, 2026
ce8865c
Add forward all-gather overlap (#5513)
wujingyue Jul 9, 2026
e86c262
Fix seq_load_balancing loss with inter-document masking and MBS > 1 (…
deepakn94 Jul 10, 2026
779c5b7
Avoid X11 master port default (#5299)
guihong-nv Jul 10, 2026
1aa880d
Fix infinite recursion in abstract tokenizer special-id property alia…
asadbekXodjayev Jul 10, 2026
2c579b4
Set Bert TE spec q/k_layernorm to None (#5687)
bbuschkaemper Jul 10, 2026
5389d14
Set is_first_microbatch when quant_recipe is configured (#5642)
yezhengmao1 Jul 10, 2026
75a2132
Assign BERT CODEOWNERS to GPT team (#5746)
Phlip79 Jul 10, 2026
48a887f
Remove use of exec_module (#5744)
jon-barker Jul 10, 2026
a28ca48
Short-circuit condition to avoid copying from GPU memory in `ChainedO…
filaretov Jul 13, 2026
3005e9c
Increase Megatron-FSDP overlap test dim to 8192 for reliable overlap …
wujingyue Jul 14, 2026
eadbaa6
Inference: Add profile endpoints to chat completions. (#5611)
sidsingh-nvidia Jul 14, 2026
3a253ac
Inference: Do not route pad/dummy tokens to any expert (#4922)
sidsingh-nvidia Jul 14, 2026
bcf4c8f
Inference: Add load aware routing to prefix caching. (#5607)
sidsingh-nvidia Jul 14, 2026
3c46452
fix(clip_grads): handle empty grads_for_norm in inf-norm and p-norm p…
Mattral Jul 14, 2026
c09022e
test(gpt): AUT-830 mark tp1_pp4_vp1_resume_torch_decoupled_lr flaky o…
svcnemo-autobot Jul 14, 2026
a79f49d
Delegate reasoning token retention to the chat template in multi-turn…
sidsingh-nvidia Jul 14, 2026
e97ac83
Various ModelOpt fixes: QAD test for CICD, use model builder config i…
AAnoosheh Jul 14, 2026
97b25d9
build: Update Transformer Engine to 2.17 (#5680)
balasaajay Jul 14, 2026
e5344ab
Mamba prefix caching fixes (#5502)
santhnm2 Jul 14, 2026
a327943
Exercise nested MFSDP CUDA graph capture (#5796)
wujingyue Jul 14, 2026
33706f9
Implement Quantile Balancing in MoE (#5349)
Mellonta Jul 15, 2026
f8c9911
Unset NCCL overrides for MFSDP v2 tests (#5794)
wujingyue Jul 15, 2026
4bf7fca
Fix MegatronFSDP root module hook dispatch (#5808)
wujingyue Jul 15, 2026
834abc1
Pin cudnn-fe and cuTeDSL version (#5812)
balasaajay Jul 15, 2026
5abab93
chore: rotate oncall schedule
github-actions[bot] Jul 15, 2026
82e9dc6
Inference: Add the nemotron_v3 reasoning parser (#5634)
sidsingh-nvidia Jul 15, 2026
ccdfa7d
Missing moe_router_dtype causes unexpected downcast in ModelOpt examp…
jinhangchoi Jul 15, 2026
5e45ccf
Avoid FSDP unit terminology in MFSDP v2 (#5793)
wujingyue Jul 15, 2026
7711346
Fix configured norm epsilon in MambaLayer (#5750)
shanhaoli Jul 15, 2026
ffbe018
[refactor] Common combined-1F1B schedule-plan base (1/4 of #4798) (#4…
Connor-XY Jul 15, 2026
ecf5547
chore(tests): AUT-851 move NCCL defaults from run_ci_test.sh to conft…
svcnemo-autobot Jul 15, 2026
8bf7365
Set num_splits to 0 for FA4 inference (#5804)
santhnm2 Jul 15, 2026
06bb399
ci: Enhance nightly/mr/weekly error reporting (#5831)
balasaajay Jul 15, 2026
3927d9b
feat(docker): Add NCCL installation script and install NCCL 2.30.4 (#…
balasaajay Jul 15, 2026
53a2dd5
IMA fix by making the copy of book keeping buffer to GPU blocking (#5…
shanmugamr1992 Jul 15, 2026
802da56
Add GPTModel to HybridModel migration guide (#5698)
Phlip79 Jul 16, 2026
38626c2
Pair frozen FSDP backward hooks (#5710)
wujingyue Jul 16, 2026
c2f0df5
Clarify NVIDIA email signing guidance (#5699)
wujingyue Jul 16, 2026
d981f66
Overlap FSDP communication with compute (#5719)
wujingyue Jul 16, 2026
7b2a97b
chore(skills): add Regent Open Plugin manifest (#5840)
ko3n1g Jul 16, 2026
2aa3645
chore(skills): remove Open Plugin manifest (superseded) (#5842)
ko3n1g Jul 16, 2026
edd4562
return prefix cache hits data from the chat completions api (#5609)
sidsingh-nvidia Jul 16, 2026
f8e1ac6
Refactor data parallel coordinator to enable modular handlers (#5550)
santhnm2 Jul 16, 2026
407ea4f
Pass device IDs to cleanup barrier (#5702)
wujingyue Jul 16, 2026
4abeecc
Fix averaging for MoE z-loss metric tracking (#3199)
Marks101 Jul 16, 2026
9a67ea0
test(mfsdp): AUT-881 mark test_overlaps_communication_and_compute fla…
svcnemo-autobot Jul 16, 2026
740c16e
Inference: Extend default cuda-graph coverage to 512 tokens (#5797)
sidsingh-nvidia Jul 16, 2026
e3fe550
Fix for sequence-level aux MoE loss being dependent on batch size (#5…
OlegSudakov Jul 17, 2026
61f3114
Allow parameterless FSDP root modules (#5711)
wujingyue Jul 17, 2026
27e312c
Fix issue where parameter groups with different min/max LRs get overr…
jstjohn Jul 17, 2026
acc7e64
fix bug where Gemma4 is not working with recompute_granularity = "ful…
warpuv Jul 17, 2026
4a1f743
Avoid extra MFSDP v2 model-weight sync memcpy (#5834)
wujingyue Jul 17, 2026
167d51d
[experimental] Add experimental/agent_compose placeholder with previe…
ISEEKYAN Jul 17, 2026
53492eb
Stabilize mfsdp_v2 overlap test by enlarging the model (#5846)
wujingyue Jul 17, 2026
dc57c2b
Print important dependencies (#5814)
gautham-kollu Jul 18, 2026
f258d4f
Test zero-CTA copy-engine all-gather (#5858)
wujingyue Jul 18, 2026
b4ad280
Route Lion through DistributedOptimizer and support single-moment che…
deepakn94 Jul 18, 2026
3219f38
Fix broken remove_sharded_tensors public API and re-enable its unit t…
huthvincent Jul 19, 2026
adfb8ec
Make the model larger and higher mb size to make reduce flakiness (#5…
gautham-kollu Jul 20, 2026
207be1f
chore(beep boop 🤖): Bump (main) (2026-07-20)
github-actions[bot] Jul 20, 2026
1c74286
Update active oncall to Phlip79 this week (#5896)
Phlip79 Jul 20, 2026
f41ec54
Overlap async scheduling phases (#5549)
lmcafee-nvidia Jul 20, 2026
e443bd3
ci: integrate nemo-ci-triage with linear issues management for gitlab…
balasaajay Jul 20, 2026
bd0872d
Reduce boilerplate around MultiStorageClient feature checks (#5269)
Randl Jul 20, 2026
cfb116f
Add NeMo waveform audio processor (data-side feature extractor) (#5570)
yqwangustc Jul 20, 2026
5a88c57
Support HSDP deferred DP-outer gradient reduction (#5743)
Achyuthan-S Jul 20, 2026
337c061
Add fully_shard_optimizer for mixed-precision FSDP (#5411)
wujingyue Jul 20, 2026
7012adb
Stabilize perf warmup (#5913)
Phlip79 Jul 21, 2026
12c05a2
Add compatibility between training CGs and CP>1 (#5894)
tdene Jul 21, 2026
368fa88
Log app_finish_time and app_train_loop_finish_time on early-exit path…
aliardaeker Jul 21, 2026
58bf14e
Test mFSDP v2 overlap with default and symmetric memory (#5859)
wujingyue Jul 21, 2026
ddaa315
fix: Harden Claude GitHub workflows (#5408)
chtruong814 Jul 22, 2026
a046281
chore: rotate oncall schedule
github-actions[bot] Jul 22, 2026
d207685
Reuse profiler helpers in mFSDP v2 symmetric memory tests (#5873)
wujingyue Jul 22, 2026
cc5b092
Reduce MimoOptimizer update-success across the world for cross-grid c…
yashaswikarnati Jul 22, 2026
96d4159
Refresh BERT H100 golden values (#5953)
Phlip79 Jul 22, 2026
890247c
fix(resharding): stabilize NVSHMEM refit copy service (#5915)
wdykas Jul 22, 2026
e4ea859
[Main] Numerical fix for FC2 expert bias scales when using `use_trans…
zhongbozhu Jul 22, 2026
b842f59
ci: Integrates the latest config-driven nemo-ci-triage Slack and Line…
balasaajay Jul 22, 2026
8d16c67
[Main] Numerical fix for moe single grouped weight with fp8 fp4 prima…
zhongbozhu Jul 22, 2026
602fad0
Inference: Do not let prompt tokens return from the engine, unless re…
sidsingh-nvidia Jul 22, 2026
52cf20d
Make LRU prefix caching eviction policy only evict child blocks (#5822)
santhnm2 Jul 23, 2026
411a5d8
Inference: Reduce mamba scratch space size by an order of magnitude. …
sidsingh-nvidia Jul 23, 2026
c5398de
Fix TE grouped MLP fused main-grad setup (#5209)
Wong4j Jul 23, 2026
bb5647a
fix(ci): AUT-957 support golden checks in merge queue (#5989)
svcnemo-autobot Jul 23, 2026
6cd6ea5
Fix formatting error in qwen3_30b_a3b config (#5978)
jon-barker Jul 23, 2026
51e915a
Batch-invariant train/inference logprob parity (#5897)
wdykas Jul 23, 2026
15d4b34
Move FSDP model weight sync to optimizer post-step (#5949)
wujingyue Jul 23, 2026
acd0e35
fix: allow mtp_num_layers=0 with overlap_moe_expert_parallel_comm (#5…
cuichenx Jul 23, 2026
4464d1c
Port Multi-Latent Attention to `HybridModel` (#4452)
janEbert Jul 23, 2026
4f72cc0
docs(skills): clarify container::lts is the older LTS PyTorch base (#…
ko3n1g Jul 23, 2026
ff60360
test(hybrid): AUT-971 quarantine Nemotron QAD functional test (#6013)
svcnemo-autobot Jul 24, 2026
81770cb
[main] add thd sequence packing dispatcher support for main (#5008)
HaochenYuan Jul 23, 2026
a69bbb0
Add encoder prefetch for heterogeneous MIMO training (#5833)
yashaswikarnati Jul 24, 2026
0a7f372
Use explicit process groups for dataloader checkpoints (#5988)
yashaswikarnati Jul 24, 2026
4f65717
[2/2] Add TileLang fused DSA kernels support with THD and CP & Clean …
HollowMan6 Jul 24, 2026
6f0db6f
Refit: local plan building, node-add cache key, and NIXL backend (#5872)
wdykas Jul 24, 2026
72411e5
Fix FSDP2 SwiGLU checkpointing. (#5714)
cspades Jul 24, 2026
2d754e6
Enable DDP communication overlap for MIMO training (#5979)
yashaswikarnati Jul 24, 2026
af3d240
rl: release G-submission gate slots on consumption instead of assembl…
lauradang Jul 24, 2026
648bc01
Inference: Optimized triton kernels to extract mamba states in prefix…
sidsingh-nvidia Jul 24, 2026
066145e
[Inference] Set different random seeds for each DP rank for generatio…
cspades Jul 24, 2026
dfe9d04
Ensure Mamba prefix cache snapshots are recorded for multi-chunk prom…
santhnm2 Jul 24, 2026
50cc27a
ci(actions): AUT-977 retry transient log artifact uploads (#6027)
svcnemo-autobot Jul 25, 2026
21fe0fe
fix: add additional error checks for flaky failures (#6029)
balasaajay Jul 25, 2026
7ee2580
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jul 25, 2026
650b783
Prevent FlashInfer sampling from running with CUDA graphs (#5791)
santhnm2 Jul 24, 2026
10c9383
fix(inference): AUT-980 disable fp8 primary weights in graph tests (#…
svcnemo-autobot Jul 25, 2026
7e6a045
chore(deps): AUT-967 stabilize Transformer Engine 2.18 upgrade (#5997)
svcnemo-autobot Jul 25, 2026
c5ff22b
[feat] Generalized Tensor Parallelism (GTP) (#4967)
fanshiqing Jul 25, 2026
cd4afff
Fix gradient-norm undercounting when using EP and TP (#5916)
philipcmonk Jul 25, 2026
7ec1bce
fix(inference): MCORE-536 report dropped prompt token lengths (#6051)
svcnemo-autobot Jul 26, 2026
543b803
[GTP][Feat] Add one-block-ahead prefetch for GTP grouped-expert weigh…
fanshiqing Jul 27, 2026
076f61f
chore(codeowners): AUT-1094 add GTP owners (#6062)
svcnemo-autobot Jul 27, 2026
10cd30c
Deprecate GPTModel in favor of HybridModel (#5911)
Phlip79 Jul 27, 2026
4b18b26
Fix CUDA graph correctness issues due to memory bugs (#5975)
jiemingz Jul 27, 2026
876016a
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jul 28, 2026
fc57daa
fix(dist-ckpt): AUT-1100 restore default strategy factories (#6065)
svcnemo-autobot Jul 27, 2026
07e958f
refactor: extract and split common logic between GDN & GDN2 (#5843)
xuantengh Jul 28, 2026
c4b4d59
Add load-time GPT-to-Hybrid checkpoint translation (#5675) (#5792)
guihong-nv Jul 28, 2026
866efa4
dist_ckpt: add --stream-ckpt-dequant to fix OOM on large FP8/MXFP8 lo…
asolergi-nv Jul 28, 2026
6f7bcd4
Populate dp process group in auto-built ProcessGroupCollection in pip…
ilml Jul 28, 2026
bacd340
fix(cuda-graphs): MB-928 align DDP initialization with capture stream…
svcnemo-autobot Jul 28, 2026
8a424b8
Enforce that the number of optimizer shards used in layout computatio…
deepakn94 Jul 28, 2026
4da9212
NCCL EP zero copy (#5735)
YangFei1990 Jul 28, 2026
7557c02
Optimize unit metadata for fused shared experts (#6053)
sraman-rgb Jul 28, 2026
053fa07
build: AUT-1117 serialize uv dependency installation (#6090)
svcnemo-autobot Jul 28, 2026
541d5ee
Extend dynamic inference asynchronous scheduling support (#5939)
lmcafee-nvidia Jul 28, 2026
c922805
Add support for non-Gym multi-turn environments (#5312)
tdene Jul 28, 2026
2576c15
Correct prefix-caching ref-count accounting (#6047)
tdene Jul 28, 2026
818b258
send pg group for distributed checkpoint validation (#6092)
wdykas Jul 28, 2026
2658704
fix(ci): AUT-1080 retry live NCCL watchdog timeouts (#6054)
svcnemo-autobot Jul 28, 2026
157c023
Prevent coordinator crash if an engine disconnects (#6025)
tdene Jul 28, 2026
bc67abd
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jul 29, 2026
8a63dd5
chore: rotate oncall schedule
github-actions[bot] Jul 29, 2026
0a447f5
Make nightly sync checks advisory (#5713)
Phlip79 Jul 29, 2026
f43ff6b
test(optimizer): MCORE-560 cover MoE gradient zero counts (#6050)
svcnemo-autobot Jul 29, 2026
b78cfd5
RL: fix logging issues from failed rollouts (#6061)
tdene Jul 29, 2026
e1b2a2f
ci: AUT-1141 skip MBridge dispatch for docs-only changes (#6116)
svcnemo-autobot Jul 29, 2026
3ff70c0
chore(ci): AUT-1135 pin GitHub Actions to commit SHAs (#6100)
svcnemo-autobot Jul 29, 2026
54045a5
chore(deps): AUT-1080 pin merged NeMo Run cancellation fix (#6110)
svcnemo-autobot Jul 29, 2026
efb3963
chore: AUT-1142 bump package versions (#6120)
svcnemo-autobot Jul 29, 2026
c7d1865
test(inference): mark prefix caching CUDA graph test flaky (#6131)
ko3n1g Jul 29, 2026
e598899
Make BertModel lm_head optional. (#5690)
bbuschkaemper Jul 29, 2026
eab6551
test(inference): MCORE-561 trust FP8 metadata in DeepSeek checkpoints…
svcnemo-autobot Jul 29, 2026
3f88509
ci: fix test notifications for weekly/release tests (#6055)
balasaajay Jul 29, 2026
8f1c99c
Allow the MInf sync API to run serve() (#6123)
tdene Jul 29, 2026
77c8772
Add disaggregated KV transfer backends (#5861)
nvcsathe Jul 29, 2026
542ff53
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jul 30, 2026
6183f9d
Fix gradient counting for muon+expert biases (#6099)
philipcmonk Jul 29, 2026
9436656
Reorganize a few docs (#6109)
Phlip79 Jul 30, 2026
8e57bb6
Perf: skip per-param copy_ dispatch in the MXFP8 param copy-back (#6094)
fanshiqing Jul 30, 2026
354b5d2
test: add async scheduling inference coverage (#6104)
lmcafee-nvidia Jul 30, 2026
ab16414
test(ci): AUT-1173 set functional tests to ft_launcher (#6128)
svcnemo-autobot Jul 30, 2026
ff6b92a
Fix --freeze-all-layers for new model builder path + unit tests (#5926)
AAnoosheh Jul 30, 2026
3fd8f33
Fix inference overflow into dummy blocks (#5950)
tdene Jul 30, 2026
175b4cc
fix(tests): prefer staged release assets (#6149)
ko3n1g Jul 30, 2026
ce8a8a7
Fix Megatron FSDP optimizer under-update (#5976)
wujingyue Jul 30, 2026
9d416ee
Rebind MFSDP v2 sharded grads once per gradient reduction (#6041)
wujingyue Jul 30, 2026
0081c84
Fix MLA dynamic inference decode flag (#4902)
cuichenx Jul 30, 2026
dc3c9ab
Normalize MFSDP data-parallel axes at fully_shard (#6067)
wujingyue Jul 30, 2026
ab20a60
Add Agent Compose package and runtime interface skeleton (#5937)
ISEEKYAN Jul 30, 2026
ec53920
Fix GTP full-iteration CUDA-graph capture regression (#6077)
prajwal1210 Jul 30, 2026
d04e7a0
Harden dynamic prefix cache allocation - mamba mostly (#6091)
wdykas Jul 30, 2026
6513e3e
Update copy-pr-bot.yaml [skip ci]
github-actions[bot] Jul 31, 2026
8543ae4
ci(cache): AUT-1188 harden container cache selection (#6160)
svcnemo-autobot Jul 30, 2026
b19b1f4
ci(nvskills): AUT-1188 fix reusable workflow startup failure (#6162)
svcnemo-autobot Jul 31, 2026
d3519d5
Add MFSDP runtime schedule design (#6031)
wujingyue Jul 31, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
1 change: 1 addition & 0 deletions .agents/skills
14 changes: 14 additions & 0 deletions .claude/settings.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
{
"hooks": {
"UserPromptSubmit": [
{
"hooks": [
{
"type": "command",
"command": "printf '{\"hookSpecificOutput\":{\"hookEventName\":\"UserPromptSubmit\",\"additionalContext\":\"MANDATORY WORKFLOW — never skip or reorder: (1) Read the artifact first (commit, file, error, PR). (2) Identify and invoke the relevant skill via the Skill tool BEFORE forming any answer or plan — even when the answer seems obvious. (3) Only then answer using the skill context. Skipping step 2 is not allowed.\"}}'"
}
]
}
]
}
}
1 change: 1 addition & 0 deletions .claude/skills
1 change: 1 addition & 0 deletions .cursorrules
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
See CLAUDE.md for all repository guidelines.
85 changes: 85 additions & 0 deletions .github/CODEOWNERS
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
megatron/core/ @NVIDIA/core-adlr @NVIDIA/core-nemo
megatron/core/tensor_parallel/generalized_tensor_parallelism.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/gtp

megatron/core/models/bert/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/gpt

megatron/core/models/common/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/gpt

megatron/core/models/gpt/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/gpt

megatron/core/models/multimodal/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/multi-modal

megatron/core/models/audio/ @NVIDIA/mcore-audio

megatron/core/models/mamba/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/hybrid-model
megatron/core/ssm/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/hybrid-model

megatron/core/models/hybrid/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/hybrid-model

megatron/core/datasets/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/datasets

megatron/core/tokenizers/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/tokenizers

megatron/core/distributed/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/data-parallelism
megatron/core/distributed/fsdp/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/megatron-fsdp

megatron/core/transformer/fsdp_dtensor_checkpoint.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/megatron-fsdp

megatron/core/dist_checkpointing/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/dist-checkpointing

megatron/core/optimizer/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/mcore-optimizer

megatron/core/optimizer/distrib_optimizer.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/dist-optimizer
megatron/core/optimizer/layer_wise_optimizer.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/dist-optimizer
megatron/core/optimizer/param_layout.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/dist-optimizer

megatron/core/optimizer/emerging_optimizers.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/mcore-emerging-optimizers
megatron/core/optimizer/muon.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/mcore-emerging-optimizers
megatron/core/optimizer/qk_clip.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/mcore-emerging-optimizers @NVIDIA/transformer

megatron/core/inference/modelopt_support @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/post-training

megatron/core/datasets/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/datasets

megatron/core/pipeline_parallel/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/pipeline-parallelism

megatron/core/transformer/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/transformer

megatron/core/transformer/moe/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/mixture-of-experts-adlr @NVIDIA/mixture-of-experts-devtech

megatron/core/inference/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/inference

megatron/inference/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/inference-interface

megatron/core/parallel_state.py @NVIDIA/core-adlr @NVIDIA/core-nemo

megatron/core/post_training/ @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/post-training

megatron/post_training/ @NVIDIA/post-training

megatron/core/transformer/cuda_graphs.py @NVIDIA/core-adlr @NVIDIA/core-nemo @NVIDIA/cuda-graphs

megatron/training/ @NVIDIA/training-adlr @NVIDIA/training-nemo
megatron/training/arguments.py

.gitlab/ @NVIDIA/ci
.github/ @NVIDIA/ci
.github/oncall_schedule.json @NVIDIA/mcore-oncall-rotation
.gitlab-ci.yml @NVIDIA/ci
docker/ @NVIDIA/ci
tests/functional_tests/python_test_utils/ @NVIDIA/ci
tests/functional_tests/shell_test_utils/ @NVIDIA/ci
tests/test_utils/recipes/ @NVIDIA/ci
tests/unit_tests/run_ci_test.sh @NVIDIA/ci

# API Backwards Compatibility Check
scripts/check_api_backwards_compatibility.py @NVIDIA/ci
scripts/README_API_COMPAT.md @NVIDIA/ci
.github/workflows/check_api_backwards_compatibility_workflow.yml @NVIDIA/ci
docs/api-backwards-compatibility-check.md @NVIDIA/ci
tests/unit_tests/test_api_backwards_compat_setup.py @NVIDIA/ci

megatron/rl/ @NVIDIA/reinforcement-learning
examples/rl/ @NVIDIA/reinforcement-learning
test/unit_tests/test_rl_utils.py @NVIDIA/reinforcement-learning
train_rl.py @NVIDIA/reinforcement-learning
29 changes: 29 additions & 0 deletions .github/ISSUE_TEMPLATE/bug_report.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
---
name: Bug report
about: Create a report to help us improve the repository or project
title: ""
labels: bug
assignees: ''

---

**Describe the bug**

A clear and concise description of what the bug is. Tag @NVIDIA/mcore-oncall
to get oncall's attention to this issue.

**Steps/Code to reproduce bug**

Please list *minimal* steps or code snippet for us to be able to reproduce the bug.

A helpful guide on on how to craft a minimal bug report http://matthewrocklin.com/blog/work/2018/02/28/minimal-bug-reports.


**Expected behavior**

A clear and concise description of what you expected to happen.


**Additional context**

Add any other context about the problem here.
2 changes: 2 additions & 0 deletions .github/ISSUE_TEMPLATE/config.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
blank_issues_enabled: false

23 changes: 23 additions & 0 deletions .github/ISSUE_TEMPLATE/feature_request.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
---
name: Feature request
about: Suggest an idea for this project
title: ""
labels: enhancement
assignees: ''

---

**Is your feature request related to a problem? Please describe.**
A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]

Tag @NVIDIA/mcore-oncall
to get oncall's attention to this issue.

**Describe the solution you'd like**
A clear and concise description of what you want to happen.

**Describe alternatives you've considered**
A clear and concise description of any alternative solutions or features you've considered.

**Additional context**
Add any other context or screenshots about the feature request here.
13 changes: 13 additions & 0 deletions .github/ISSUE_TEMPLATE/question.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
---
name: QUESTION
about: Ask a question about Megatron-LM that is not a bug, regression or enhancement
request
title: "[QUESTION]"
labels: ''
assignees: ''

---

**Your question**
Ask a clear and concise question about Megatron-LM. Tag @NVIDIA/mcore-oncall
to get oncall's attention to this issue.
40 changes: 40 additions & 0 deletions .github/ISSUE_TEMPLATE/regression.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
---
name: REGRESSION
about: Report a regression in speed or accuracy due to a Megatron-LM update
title: "[REGRESSION]"
labels: ''
assignees: ''

---

**Describe the regression**
A clear and concise description of what the regression is. Tag @NVIDIA/mcore-oncall
to get oncall's attention to this issue.

**To Reproduce**
Steps to reproduce the behavior. The easier it is to reproduce the faster it will get maintainer attention.

**Previous performance**
What speed or accuracy did you previously see.

**New performance**
What speed or accuracy do you see after the update.

**Stack trace/logs**
If applicable, add the stack trace or logs related to the regression.

**Environment (please complete the following information):**
- Previous Megatron-LM commit ID
- New Megatron-LM commit ID
- Previous PyTorch version
- New PyTorch version
- Previous CUDA version
- New CUDA version
- Previous NCCL version
- New NCCL version

**Proposed fix**
If you have a proposal for how to fix the issue state it here or link to a PR.

**Additional context**
Add any other context about the problem here.
Loading