😯
SIU
Incoming CS PhD @stanfordnlp Building continuously self-improving foundation model agents
Highlights
Pinned Loading
-
spade-rl/spade
spade-rl/spade PublicSPADE: Self-Play in Adaptive Synthetic Executable Environments
-
spiral-rl/spiral
spiral-rl/spiral PublicSPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
-
deepseek-ai/DeepSeek-V2
deepseek-ai/DeepSeek-V2 PublicDeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
-
deepseek-ai/DeepSeek-VL
deepseek-ai/DeepSeek-VL PublicDeepSeek-VL: Towards Real-World Vision-Language Understanding
-
metaopt/torchopt
metaopt/torchopt PublicTorchOpt is an efficient library for differentiable optimization built upon PyTorch.
-
waterhorse1/Natural-language-RL
waterhorse1/Natural-language-RL PublicNatural Language Reinforcement Learning
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.





