Tom Lu
eigentom
AI & ML interests
MLLM, Reinforcement Learning, Agentic RL
Recent Activity
published a model 1 day ago
eigentom/qwen35-4b-dci-rl-rewardv2-step10 published a dataset 7 days ago
DCI-Agent/dci-experiment-corpus upvoted a paper 15 days ago
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation