I trained a small transformer in 1.5hrs and it beats many LLMs
Trending on Hacker News: I trained a small transformer in 1.5hrs and it beats many LLMs (510 points / 142 comments, via mvakde.github.io)
Opening excerpt
I trained a small transformer from scratch in 1.5hrs on a 5090 Beats many LLMs, and scores the same as TRM/HRM
This is an upgrade to my previous model Faster, better, cheaper and still open source.
Many ppl thought the prev result was impossible. It got attention from top researchers and went viral on X. Eg: Discussions by Lucas Beyer , Jeremy Howard , Rohan Anil , and comments by many others.
I think sample efficiency is the most important problem in AI today and I want to solve it.
(Excerpted from the original; full article via the source link below.)
This story hit the Hacker News front page today (510 points / 142 comments, via mvakde.github.io). Our Tech Radar aggregates daily signals on AI engineering, backend architecture and DevOps — browse the related services and further reading below, or get in touch with our team.
Source: Hacker News