Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
|
Tejush's comments
login
Tejush
3 days ago
|
parent
|
context
|
next
[–]
| on:
DeepSeek v4.1 Flash is now available for internal ...
based
reply
Tejush
46 days ago
|
parent
|
context
|
prev
[–]
| on:
Kimi-K3 Technical Report [pdf]
Any guess on the pre-training tokens/flops they consumed?
lostmsu
46 days ago
|
parent
[–]
Technical report has a graph vs Kimi K2 with 1e21 FLOPs (but they don't claim that's the entirety of pretraining)
JacobAsmuth
46 days ago
|
root
|
parent
[–]
1e21 flops is hilariously wrong. for reference the llama 3 8B model (
https://arxiv.org/pdf/2407.21783
) used 10 times that many flops. This model is 350x bigger in total params and 12x bigger in active params and was trained on 3x the data.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
reply