We've always gotten a ton of positive comments from people with ADHD who've used our HTML videos to learn to code. Despite it not being in our target group at all when we started out.
I gave it a fair shake (10 or so).
I'll say this... it does the job of "I can't spend 45 minutes on a journal article that's already over my head, or get through 20 pages of blog slop, so please give me the summary version of this".
The select comment/conversation slide is a nice touch, though it's sure making some choices there.
Main problem seems to be that the sameness of the format (it's clearly following a strict template) and voice do make it feel very repetitive, real fast. I'm not sure how you get consistent results AND solve for that, though.
These AI explainers are taking over. I tried getting this working back in May. The models were okay but not quite there and it was a ton of effort. Opus 5.5 seems like the tipping point.
That video in the README is amazing! Will definitely dig into this project.
We're not using Opus 5.5 btw, as it would be too expensive. Our goals is to eventually get the cost down to 1 cent for a 1 minute video, and with 1 second delay from submit to playback. For that, we need dirt cheap and lightning fast models.
honestly it's horrible compared to some of the Opus 5.5 ones. But yeah, at the latency/speed you're looking for it's a different ballpark. Let Opus write the style and reference screens/animations, let your small fast model assemble it from parts.
This is great - many of us who are technical but not working in hardcore tech probably scroll by quickly on stuff we have never heard of. Getting a quick synopsis like this may drive more traffic and understanding.
This is actually quite impressive from an engineer viewpoint. I just have the feeling that the videos quickly become very monotonous and rather boring due to the monotonous AI voices. If somehow you could bring dynamic variation in these videos that would be fantastic.
All of what you said is great, and I especially detest the proliferation of slop videos on YouTube, where it makes no sense to drown out the ample supply of such personalities and insights with AI voices summarizing wikipedia or Reddit threads over AI imagery slideshows.
But given how good a job this seemed to do at giving me more than just headlines, I'd love to have maybe even just an audio podcast feed of the top 10 stories like this compiled a few times a day. I would listen while I'm doing things when reading isn't practical.
tl;dr agree that we don't need this to replace reading, but I see that it can be a useful tool.
Yeah I do see your point. And for some reason making it into audio does seem less offensive to me. Maybe it’s because it at least transforms it into something I can do while doing something else. Whereas a video is basically the same mode of operation: staring at a screen.
I’m still basically against this though. I consider it slop.
I just had opus make me a tts mobile app that grabs the top 20 HN stories and the top 5 comments for each and read them out to me (using one of the higher quality built in TTS voices on my pixel). The summarization is the worst part of this imo, so pure TTS wins
Presumably you also like that they are explaining things right? I mean that seems like the more critical step. Otherwise if it's just Marquise Brownlee, you could just watch Marquise Brownlee say the same sequence of random words for 3 minutes a few times per day.
I think this is exactly the opposite of what AI should be used for. It is going to make people dumb.
Watching a video instead of engaging your own brain makes you feel like you learned something without actually trying and I would bet it works much less well.
If someone is there adding something beyond what is already there - ie someone like Maruqise. Then it makes sense for them to be there.
If not it is just brain rot.
People already have issues with concentration. Allowing them to further allow that muscle to atrophy will not be good.
Thanks! And I agree. We're getting tons of requests for more nuance in voice selection from our users. You're currently able to say i.e. "Australian English female" in your prompts today, but you should also ideally be able to describe the voice characteristic (i.e. like a funny grandpa, engaged news reporter).
What Gemini 3.8 Flash TTS is doing with generative voice design in this area super interesting.
This feels like when a decent book is made into a movie, and then the movie does well so someone will take the movie and summarize the plot in a YouTube video. TLDR: this is TLDR for TLDR. Just read more from the source.
Yeah I feel like this would be better as summaries of textbook chapters or information rich sources, a lot of HN posts are already basically text summaries of a complex topic that it doesn't really make sense to further summarize them.
To me, the use case is to go from an often opaque headline ("Using Nix and containerd with Jev on macOS") to a paragraph which hopefully gives some strong hints as to what those things are and why it's interesting, because the articles often are written for an audience already deeply into all the topics. Because often a few of those terms I might have no clue about and thus scroll on by, but with a brief "why you should care" explanation, I might realize it's actually something cool.
We are not particularly pleased with the state of our visualizations, so I totally agree. If you have any examples of how you'd like to see the visualizations, please share them!
Thanks! Several people on our team are actually exactly like you: they don't really use videos themselves for learning, but see the utility for others, and enjoy the technical challenges around building a high-performant video format.
Indeed. In much the same vein, I simultaneously accept that:
1. I loathe video with the white-hot glow of a thousand angry suns, and loathe LLM slop-digests of things more still.
2. I spend tremendous amounts of time in situations where I can listen but not read, whether driving or on my bike, and this is a legitimately useful way to catch up on HN in those situations.
I would stick with this. There's huge potential here. I'm not going to lie, I did a video of this, and it was pretty slop. Some of it didn't even make sense and was super irrelevant. That being said, refined, there's really actually some major potential with this idea.
I don't dislike that it's video (because I know people who won't read can't be "made to read" by there not being "nice enough explainer videos"), but this is the kind of stuff that bothers me in both video and text, and greatly. I don't want to rant about it here because it's not specific to your product at all, it's not even specific to LLM, the internet and especially youtube is full of stuff that neither speakers/authors nor audience seem to ever parse. And if you're just downstream of big models you probably can't really influence that. But since you are NIH enjoyers, maybe you could train your own at some point? I don't want to consume such videos, but I do want those who do to have nicer ones, because that'd be good for me, too.
I just tried it out - it's pretty cool. The postscript was actually noted in the generated video about this topic, which I find amusing.
I have to say, if HN were to add an AI generated paragraph summary about each link at the top of each comments page, it would probably go a long way to improving the discussions about a lot of topics. The number of people who comment based on the title of a link alone is pretty high, and I'd bet the vast majority of commenters only skim the linked pages anyways. Might as well get everyone on the same page with a quick overview.
For most articles these days, we'd have to first have a separate bot get the archive.is link for it!
But I'm skeptical that people would embrace it. It's more about the optics, and it also relates to the old principle of "Read the article -- don't start opining based only on the headline." Many would probably say that you shouldn't offer an opinion if you've only read an AI distillation.
It could be wrong - or more likely, the article acknowledges likely objections and refutes them well, but that part didn't make it into the summary, so everyone starts raising very un-insightful points as though the author was oblivious to them.
This looks a good idea to let people choose their preferred medium. I am not into video either, but I turned one of my latest blog article into a video using Claude, since I thought it would be “easy.” I asked it to build the video by screencasting a browser presenting the interactive diagrams of the article and add some titles. It build some kind of recording app for me. It still took me many hours, but I think I would be faster if I redid it. I would pay to have it done faster, but I want to keep control of the text, not have a summary.
You can use MCP to generate an Explainer video next time if you want. It's currently free, as the tokens for directing and scripting it is offloaded to your agent. And then we cover the TTS for the time being.
I like the concept and these explainers. It sounds like this is using a HTML based format to be similar to like a Flash / Shockwave animation to keep data size down.
I wonder what the challenge would be of making these videos more interactive would be? Also I do wonder if there is any research happening in making AI explainers more trustable.
Great. Wordpress plugin would work. Get publishers to embed these videos at the top of their own articles. Allow them to tweak/fix/fine-tune them. Accept micropayments and on-demand generation. Readers can pay for better videos or "deep-dives."
This will work for young readers too, illiterates, sensory-impaired and foreigners. All four!!
Nice concept, but why not generate one video per post rather than re-generating the video for each click per user? The result is going to be the same, and it will be significantly cheaper to run.
Btw, we are hiring a new developer to join us in Oslo, Norway. We pay well and give a lot of freedom wrt working from home or at the office (hybrid solution). If getting HTML-based video generation down to 1 cent for 1 minute with 1 second delay, then please email me: per@scrimba.com.
We currently spend ~$0.4 on a video (without images) and a have a few seconds of delay before playback. So it's a matter of turning every stone in order to optimize it further. We're going to add Jev asap, as it's an obvious improvement along both axes.
That was one of the challenges when building this. We ended up using Firecrawl. But even so there are a few pages we aren't able to read. In which case we use the Algolia API to get the HN comments and generate the video about the HN reaction on the article instead.
DANmode | 4 hours ago
Kidding, I already sent it to a friend with ADHD who has been struggling to remain anchored to the tech world in any way besides Shorts.
But, culturally, I definitely hate the trend it implies!
Neat. Thank you for sharing.
nxc18 | 3 hours ago
I watched the videos and confirmed they are worse than brainrot.
[OP] mrborgen | 55 minutes ago
We've always gotten a ton of positive comments from people with ADHD who've used our HTML videos to learn to code. Despite it not being in our target group at all when we started out.
Hope your friend finds it helpful!
dd8601fn | 33 minutes ago
The select comment/conversation slide is a nice touch, though it's sure making some choices there.
Main problem seems to be that the sameness of the format (it's clearly following a strict template) and voice do make it feel very repetitive, real fast. I'm not sure how you get consistent results AND solve for that, though.
alentodorov | 4 hours ago
[OP] mrborgen | an hour ago
babu_mick | 4 hours ago
mgxplyr | 4 hours ago
harvey9 | 4 hours ago
nonethewiser | 3 hours ago
This is pretty crazy. It's not hard to imagine something like Reddit deploying this as a first party feature.
simbas | 4 hours ago
I'm going to be looking to reproduce this.
[OP] mrborgen | 2 hours ago
scosman | 4 hours ago
I made an OSS framework for these for when you want to go beyond one-shoting it: https://github.com/scosman/videowright
- Voiceovers: aligns animations to the voiceover, can generate voiceover with elevenlabs, or will transcribe and timestamp a real voiceover
- can reorder scenes both in code, and using ffmpeg for audio.
- interactive controls during authoring, can ask for micro edits or re-builds
- MP4 export/encoder
- Generates a video from a prompt (obvs)
[OP] mrborgen | 2 hours ago
We're not using Opus 5.5 btw, as it would be too expensive. Our goals is to eventually get the cost down to 1 cent for a 1 minute video, and with 1 second delay from submit to playback. For that, we need dirt cheap and lightning fast models.
scosman | 2 hours ago
Vaslo | 4 hours ago
Thanks!
dverlaeckt80 | 4 hours ago
jonplackett | 2 hours ago
Like Marquise Brownlee just has opinions I care about and I watch his videos for this reason.
If a video just explains something I could just read all I’m getting is a layer of obfuscation.
xp84 | 2 hours ago
But given how good a job this seemed to do at giving me more than just headlines, I'd love to have maybe even just an audio podcast feed of the top 10 stories like this compiled a few times a day. I would listen while I'm doing things when reading isn't practical.
tl;dr agree that we don't need this to replace reading, but I see that it can be a useful tool.
jonplackett | 2 hours ago
I’m still basically against this though. I consider it slop.
TomGarden | an hour ago
nonethewiser | 2 hours ago
jonplackett | 2 hours ago
I think this is exactly the opposite of what AI should be used for. It is going to make people dumb.
Watching a video instead of engaging your own brain makes you feel like you learned something without actually trying and I would bet it works much less well.
If someone is there adding something beyond what is already there - ie someone like Maruqise. Then it makes sense for them to be there.
If not it is just brain rot.
People already have issues with concentration. Allowing them to further allow that muscle to atrophy will not be good.
itomato | 2 hours ago
[OP] mrborgen | 2 hours ago
What Gemini 3.8 Flash TTS is doing with generative voice design in this area super interesting.
goshx | 4 hours ago
daok | 4 hours ago
1970-01-01 | 3 hours ago
redhed | 2 hours ago
xp84 | 2 hours ago
j45 | 3 hours ago
[OP] mrborgen | 2 hours ago
fishtoaster | 3 hours ago
1. I hate everything about this because I vastly prefer text over video for the same content, especially AI generated video
2. There are a lot of people for whom video is their preferred medium and so this will be valuable to them.
It's a technically cool project and your cost-per-video is impressively low. Best of luck!
jonplackett | 3 hours ago
kekebo | 2 hours ago
[OP] mrborgen | 2 hours ago
abalashov | an hour ago
1. I loathe video with the white-hot glow of a thousand angry suns, and loathe LLM slop-digests of things more still.
2. I spend tremendous amounts of time in situations where I can listen but not read, whether driving or on my bike, and this is a legitimately useful way to catch up on HN in those situations.
mlaretallack | 3 hours ago
ProofHouse | 3 hours ago
[OP] mrborgen | 54 minutes ago
Think there's a ton of potential here too. So we're staying on course.
customguy | 3 hours ago
"a compressor uses three main organs" @ 2:97
I don't dislike that it's video (because I know people who won't read can't be "made to read" by there not being "nice enough explainer videos"), but this is the kind of stuff that bothers me in both video and text, and greatly. I don't want to rant about it here because it's not specific to your product at all, it's not even specific to LLM, the internet and especially youtube is full of stuff that neither speakers/authors nor audience seem to ever parse. And if you're just downstream of big models you probably can't really influence that. But since you are NIH enjoyers, maybe you could train your own at some point? I don't want to consume such videos, but I do want those who do to have nicer ones, because that'd be good for me, too.
russellbeattie | 2 hours ago
I have to say, if HN were to add an AI generated paragraph summary about each link at the top of each comments page, it would probably go a long way to improving the discussions about a lot of topics. The number of people who comment based on the title of a link alone is pretty high, and I'd bet the vast majority of commenters only skim the linked pages anyways. Might as well get everyone on the same page with a quick overview.
xp84 | 2 hours ago
But I'm skeptical that people would embrace it. It's more about the optics, and it also relates to the old principle of "Read the article -- don't start opining based only on the headline." Many would probably say that you shouldn't offer an opinion if you've only read an AI distillation.
It could be wrong - or more likely, the article acknowledges likely objections and refutes them well, but that part didn't make it into the summary, so everyone starts raising very un-insightful points as though the author was oblivious to them.
vbernat | 2 hours ago
The article: https://vincent.bernat.ch/en/blog/2026-spanning-tree-video
The tool built by Claude to make the video: https://github.com/vincentbernat/vincent.bernat.ch/tree/2808...
[OP] mrborgen | an hour ago
You can use MCP to generate an Explainer video next time if you want. It's currently free, as the tokens for directing and scripting it is offloaded to your agent. And then we cover the TTS for the time being.
kristopolous | 2 hours ago
[OP] mrborgen | 2 hours ago
kristopolous | an hour ago
fragmede | an hour ago
eleventen | 2 hours ago
[OP] mrborgen | an hour ago
bobajeff | 2 hours ago
I wonder what the challenge would be of making these videos more interactive would be? Also I do wonder if there is any research happening in making AI explainers more trustable.
adrianwaj | 2 hours ago
This will work for young readers too, illiterates, sensory-impaired and foreigners. All four!!
nonethewiser | 2 hours ago
jameshart | 2 hours ago
samplifier | 2 hours ago
wewewedxfgdf | 2 hours ago
[OP] mrborgen | an hour ago
popey | 2 hours ago
[OP] mrborgen | an hour ago
george_max | an hour ago
[OP] mrborgen | an hour ago
"We create them on-the-fly the first time someone clicks on a link."
Though I realise it could have been said more explicitly. But the second time someone clicks the link we re-use the same explainer.
ElijahLynn | an hour ago
It's definitely a form of TLDR, but for watching?!
This feels like some kind of parallel to Jev in that. If it is super fast, what other use cases may be unlocked it, as you suggest.
Also, I love the simple, easy to remember domain name, HN.watch!
[OP] mrborgen | an hour ago
[OP] mrborgen | an hour ago
wewewedxfgdf | an hour ago
How would you do that?
[OP] mrborgen | an hour ago
OutOfHere | an hour ago
[OP] mrborgen | an hour ago
OutOfHere | an hour ago
wewewedxfgdf | 16 minutes ago
Seems like over engineering - why not just make JavaScript/HTML?