Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
nolist_policy
7 months ago
|
parent
|
context
|
favorite
| on:
Normalizing Flows Are Capable Generative Models
Gemma 3n E4B runs at 35tk/s prompt processing and 7-8 tk/s decode on my last last last gen flagship Android.
ivape
7 months ago
[–]
I doubt this. What kind of t/s are you getting once your context window is reasonably saturated? Probably slows down to a crawl making it not good enough yet (the hardware that is).
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: