Out of curiosity and semi unrelated — why do so many of these projects with customized encoder-decoder setups use earlier Qwen versions like 2.5 and 3 and not the smallest 3.5? Purely the few 100m params, or something else in the latter’s arch or pretraining?
> They were building in stealth for 2 years, I was building in stealth for 2 hours…
> Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe.
By steeve 3 hours ago
No question OSS is amazing, but this video is a satire at best.
It doesn't take much attention to see the results on right vs. left side are significantly different.
Jev is not interesting if it's not "smart", a 1B param model is most definitely not smart.
By vrc a minute ago
By steeve 3 hours ago
By flockonus 3 hours ago
By mmastrac an hour ago
By tomrod 2 hours ago
By rochansinha 6 hours ago
By DavCreator 5 hours ago
By _superposition_ 4 hours ago
By looksjjhg 2 hours ago