-
Notifications
You must be signed in to change notification settings - Fork 1.6k
Pull requests: modelscope/ms-swift
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
docs: document OrcaRouter as an OpenAI-compatible sampling provider
#9995
opened Aug 27, 2026 by
nissrin2020ali-ux
Loading…
1 of 4 tasks
fix(npu): reuse fused AdamW group caches
#9984
opened Aug 25, 2026 by
addsubmuldiv
Collaborator
•
Draft
4 tasks
feat(fsdp2): load model on meta device for non-rank0 ranks (0 CPU RAM per worker)
#9982
opened Aug 24, 2026 by
cben484
Loading…
[megatron] Fix router replay CP partition and expert index overflow
#9981
opened Aug 24, 2026 by
taking-lying-flat
Contributor
Loading…
fix(fsdp): set ACCELERATE_USE_FSDP so device_map and cpu_ram_efficient_loading take effect
#9980
opened Aug 24, 2026 by
cben484
Loading…
3 of 7 tasks
fix: align channel loss with sequence parallel labels
#9977
opened Aug 24, 2026 by
taking-lying-flat
Contributor
Loading…
1 of 4 tasks
Fix Megatron-FSDP DTensor checkpoint save/load compatibility
#9976
opened Aug 24, 2026 by
addsubmuldiv
Collaborator
Loading…
1 of 4 tasks
perf(megatron): defer loss all-reduce to log time
#9966
opened Aug 22, 2026 by
gakkiri
Contributor
Loading…
1 task done
Fix zero multimodal learning rates in Megatron
#9948
opened Aug 19, 2026 by
taking-lying-flat
Contributor
Loading…
Fix Megatron GRPO vocab-parallel log-prob gradients
#9947
opened Aug 19, 2026 by
taking-lying-flat
Contributor
Loading…
Add --dataloader_multiprocessing_context to work around Python 3.14 incompatibility
#9941
opened Aug 18, 2026 by
sliedes
Loading…
1 of 4 tasks
[Megatron] Preserve RNG state across checkpoint resume
#9935
opened Aug 17, 2026 by
taking-lying-flat
Contributor
Loading…
[Megatron] Fix reference adapter loading for LoRA RLHF
#9934
opened Aug 17, 2026 by
taking-lying-flat
Contributor
Loading…
Fix reward model margin broadcasting and alignment
#9927
opened Aug 16, 2026 by
taking-lying-flat
Contributor
Loading…
Fix DPO IPO log-prob normalization
#9925
opened Aug 16, 2026 by
taking-lying-flat
Contributor
Loading…
Fix BatchSamplerShard tail sampling
#9907
opened Aug 14, 2026 by
taking-lying-flat
Contributor
Loading…
[Train] Reduce padding-free embedding output memory
#9893
opened Aug 11, 2026 by
taking-lying-flat
Contributor
Loading…
fix(deploy): handle reasoning-only stream chunks
#9881
opened Aug 11, 2026 by
taking-lying-flat
Contributor
Loading…
1 of 4 tasks
Fix IterablePackingDataset workers to use spawn
#9876
opened Aug 9, 2026 by
zupengwang
Loading…
1 of 4 tasks
[Feature] Add --streaming_shard: split the streaming dataset across data-parallel ranks
#9860
opened Aug 5, 2026 by
Rapisurazurite
Loading…
1 of 4 tasks
Previous Next
ProTip!
Follow long discussions with comments:>50.