- Alexander Wei / OpenAI Researcher2551.3/ We’ve come a long way since last summer.›LLM score 70 · 12 months ago
- Sheryl Hsu / OpenAI Researcher
- Sheryl Hsu / OpenAI Researcher
- Juntang Zhuang / MTS at xAI (pre-training lead)2554.It’s extremely fun though tough to train the first natively multimodal model ever in xAI.›LLM score 92 · 12 months ago
- Geoffrey Hinton2555.A major cut to the funding of the National Science Foundation would be very bad for the future of the US.›LLM score 20 · 12 months ago
- Ted Sanders / OpenAI Researcher2556.a cool thing you get to see building AI products: ›LLM score 75 · 12 months ago
- Ted Sanders / OpenAI Researcher2557.GPT-5 is here! it's way better at coding - not just in pointless evals, but real usage.›LLM score 70 · 12 months ago
- Jeremy Bernstein / Thinking Machines Researcher2558.I had wondered why there was no official Dion implementation by the authors...›LLM score 75 · 12 months ago
- Sally Zhu / Researcher at Flapping Airplanes2559.Interesting that you need the teacher + student to share the same base model; reminds me of linear mode connectivity stuff.›LLM score 85 · about 1 year ago
- Tri Dao / Chief Scientist at Together2560.Hierarchical layout is super elegant.›LLM score 85 · about 1 year ago
- Jason Wei / AI Researcher at Meta
- Jason Wei / AI Researcher at Meta2562.New blog post about asymmetry of verification and "verifier's law": https://t.co/bvS8HrX1jP›LLM score 80 · about 1 year ago
- Jakub Pachocki / OpenAI Chief Scientist2563.I am extremely excited about the potential of chain-of-thought faithfulness & interpretability.›LLM score 80 · about 1 year ago
- Lilian Weng / Thinking Machines Cofounder
- Tri Dao / Chief Scientist at Together2565.I played w it for 1h. Went through my usual prompts (math derivations, floating point optimizations, …).›LLM score 35 · about 1 year ago
- Tri Dao / Chief Scientist at Together2566.@RaghuGanti @cHHillee Oh you’d want to use warp reduction if the whole row fits into 1 warp.›LLM score 80 · about 1 year ago
- Tri Dao / Chief Scientist at Together2567.They’ve finally done it. They got rid of tokenizers! https://t.co/x4CXHdCw0WLLM score 60 · about 1 year ago
- Tri Dao / Chief Scientist at Together
- Tri Dao / Chief Scientist at Together2569.Getting mem-bound kernels to speed-of-light isn't a dark art, it's just about getting the a couple of details right.›LLM score 85 · about 1 year ago
- Tri Dao / Chief Scientist at Together2570.Albert articulates really well the trade offs between transformers and SSMs.›LLM score 80 · about 1 year ago
- Tri Dao / Chief Scientist at Together
- Shuchao Bi / Meta Researcher2572.
- Yang Chen / Nvidia Research Scientist2573.The first thing we did was to make sure the eval setup is correct!›LLM score 92 · about 1 year ago
- Yang Chen / Nvidia Research Scientist2574.📢We conduct a systematic study to demystify the synergy between SFT and RL for reasoning models.›LLM score 92 · about 1 year ago
- Geoffrey Hinton
- Yang Chen / Nvidia Research Scientist2576.Does RL incentive reasoning capability over the starting SFT model? ›LLM score 92 · about 1 year ago
- Geoffrey Hinton2577.I just watched a great compilation of various people's views about what is coming:›LLM score 10 · about 1 year ago
- Ludwig Schmidt / Anthropic MTS
- Geoffrey Hinton2579.AGI is the most important and potentially dangerous technology of our time.›LLM score 70 · over 1 year ago
- Geoffrey Hinton