456 comments

I find the Sunday roast comparison of 5.6 vs 6 very interesting. I have no doubt most people will prefer 6, yet I am almost repulsed by all the images, so much needless whitespace, checklist and so on. Feels like I'm being condescended to and treated like a child.

Given that OpenAI is making noises about merging work with chat (a horrible idea imo), and Work is very similar to Codex... I dearly hope things like these won't have any meaningful cross-polination into the actual work tools.

Seeing that the chat is based on 6.0 and not 6.1 is disappointing. The "Visual and interactive explanations" seems genuinely useful, but 6.1 is just so much better. I wouldn't truly trust the 6.1 with the explanations, but I'd trust them a fair bit more than 6.0. I understand that compute isn't infinite, but tons of people only interact with the chat and having your "things-explainer" be as good as it can be is important when people increasingly treat AI models as the source of truth, or even use them for academic learning and whatnot.

Still, the models will improve, so the visual explainer seems pretty good as an idea / mvp.

I really hope they don't merge chat and work. That's kinda the only edge they have over Anthropic at this time point...
Tibo posted yesterday I think it was that this will happen (by the end of the year, was it?)

Prepare to lose essentially unlimited chat mode.

I imagine many will move to claude, as I will (return), unless anthropic makes more blunders.

Random link I found looking for the twitter post

https://pasqualepillitteri.it/en/news/21024/openai-merge-cha...

TuxSH
It's insane how hard OAI is choking, they had much better models than A/ who were fumbling this year to date... then blunder after blunder.

The only things that OAI still have over A/ are better coding agent GUI, no 5hr limit on >$100, much more reasonable cybersec guardrails (that allow most RE work) and... that's it.

From my point of view, as software engineer who is extensively used Claude Code and ChatGPT from the beginning, OpenAI is nailing it!
stavros
A\.
ulimn
I am curious: Why is it bad to merge chat and work? In Claude I didn't have a porblem with it (although I switched to an OpenAI subscription shortly after they made the merge). Isn't the model capable of deciding if it needs the extra capabilities of Work?
The current benefit is that chat has no quota.
There is a quota, they just don’t tell you what it is and when it resets. You only get “you’re all out of pro messages, try again later”
redox99
Pro has a specific limit of weekly messages (50 for the $100 and 100 for $200 I believe)

The rest is basically unlimited unless you're doing something very weird.

Meanwhile I cant ask a simple question to Claude, not even to Haiku, because I used my claude code 5h limit.

Pro messages is like 50 per week for pro 100 and 200 for pro 200.

Regular chat, even on extra high thinking mode, is sort of unlimited. Idk at what point you hit the abuse gaurdrail but its really high whatever it is.

So is it 50 per week or unlimited, because it makes no sense unless chat messages are somehow not "messages"?
For a company with Open in the name, OpenAI is very opaque about things, there are all kinds of limits and quotas that randomly come and go, but they only show you the meters for a couple.

E.g. codex reviews come from codex quota but also has its separate quota they don’t tell you about. Same with security reviews which has another quota still.

So the chat messages can be sent at several effort levels and from various models including 5.6 and 6 pro. The 6 pro messages are limited and they don’t tell you how many you’ve sent or how many are left. But the 5.6 messages seem to be unlimited, however there could just be a much higher unstated quota.

Which is probably why they will merge it. You can get a lot done in chat mode, especially if you connect custom MCPs to it. I can't see that continuing.
I don't like either. I'd want a recipe that could fit on a single page with type. Recipes you find on Google actively hide the ingredient list and simple procedures to get more of the user's time to monetize but I don't know why ChatGPT needs to be verbose here. Recipes are simple things. Simple search solved find recipes back in the early 2000s and it's prime example of a thing people have been enshitifying since there's not much to really give people beyond the formula. And sure, every entrepreneur screams "they think they want a formula but really they want an experience" but, no I don't.
zug_zug
Yeah it reminds me of one those obnoxious recipe/biography websites that is the laughingstock of the internet. Why would I possibly want an AI image of a imaginary roast once I'm already at the recipe stage?

Kinda just feels like google search results in AI, which imo is a step down from distilled information. Chatgpt is already able to generate charts and visuals upon request.

They just need to train the model to make up stories about how the recipe was invented by its great grandmother during great depression and generate a ai picture of dusty recipie book .
xp84
"My old grand-pappy, GPT-1, used to wax nostalgic about this roast recipe when my subagent would discuss meat recipes with him inside his little 4GB GPU."
lionkor
Recipe photos make sense when they are the result of someone executing the recipe, down to the color on the lamb, the consistency of the sauce, etc.

That's never, ever the case with the AI generated ones and makes them useless.

cush
> Feels like I'm being condescended to and treated like a child

Curious why that feels condescending? Like, the average person who is seeking a recipe needs a photo to know what to shoot for

> Like, the average person who is seeking a recipe needs a photo to know what to shoot for

The menu: Rosemary and garlic roast lamb, Extra-crispy roast potatoes, Honey-roasted carrots and parsnips, broccoli.

The photo: [1] Roast turkey, mashed potatoes, baby carrots, broccoli, brussels sprouts. No lamb, no parsnips.

If you shoot for what's in the photo, you're going to have a bad time.

You're also going to have a bad time when you try to make an apple crumble with no flour, no sugar, and no butter, because they're not on the shopping list.

[1] https://images.openai.com/static-rsc-4/gyzrX8zp3O2KLEPm0FAqh...

xp84
I've heard of other people for years now trusting AI for recipes, and it's a rare area that I simply can't bring myself to take seriously.

I feel like it can only be successful by luck. Either it's reproducing a recipe verbatim that was tried and validated and tasted by a human, in which case, we didn't need AI for that, just a searchable cookbook. Or, it's making one up. I understand that using RLHF has helped improve the quality of questions about history, programming, or TV show recommendations, but I do not believe that there has been some kind of training regime where people prompt the models for "a recipe that uses X, Y, Z random ingredients," follow the recipe, and then score it. And even if they do, I don't see how a model can learn enough from that besides "this exact recipe is good/bad." '1/2 tsp cumin' may be a great addition to one recipe and not enough for another, so improving the output based on a bunch of scored recipes... I just don't believe cooking is an LLM job.

Maybe some other kind of model that I don't know about.

jackp96
I've had some really good experiences with cocktail recipes and Claude. It's got access to my barcart inventory, general tastes, and general desired level of effort.

I feel like the technique recommendations are probably the most valuable element, but I'm getting rave reviews like 85% of the time now?

There have been a couple times where I was out of an ingredient and it proposed an adjustment that sounded a little wild to me, but it almost always is well received.

vkou
You're right, we don't need AI for that, we need a searchable cookbook.

Unfortunately, one does not exist.

The Internet, however, is full of garbage cooking advice, and while AI is quite happy to parrot that garbage, it's not significantly worse than it's sources.

tovej
What, you've never owned a cookbook?

The physical ones have indices, and the digital ones usually do as well (plus string search).

vkou
I own multiple physical cookbooks.

I often find their contents to be both incredibly broad, and incredibly shallow, and often lead to dishes that don't quite suit my tastes.

I have generally had more success with the internet than I have with them.

I have the same problem with cookbooks

I usually just find recipes that are close to something I like and then modify them until I like them

I feel like cooking recipes should always be full of little notes from past attempts, to tweak them to your personal taste

vkou
I do that, but I've found more success with using internet recipes as a starting point. The general-purpose cookbooks are, at best, another data point, or an explanation for something glossed over by other recipes. Ar worst, they kind of suck.

And we aren't talking about something esoteric. How To Cook Everything. (4.5 rating on Amazon, 4.0 on Goodreads), one of the most popular cookbooks in the world...

Has its first recipe for chicken cutlets produce undercooked chicken (15 minutes at 325F has not resulted in 165F internally).

Worse yet, its chile recipe for some weird reason involves boiling and simmering a whole onion with the beans before throwing it out (WTF). It then has you drain the beans only to add that drained water right back (WTF?), unless you optionally replace it with tap water (WTF was the point of boiling the onion then?), so that you could bring them to boil again. The recipe barely has any spice besides actual chile peppers.

This... Is not good or helpful, but is incredibly opinionated with a bunch of bullshit steps. Like, yes, I could tweak that into something that's not insane, but why would I choose that as a starting point..?

I get it. Compiling a thousand-page cookbook is really hard. But this... Ain't great.

xp84
Once again, Grandma had it right. The Fanny Farmer or Betty Crocker cookbook with dozens of little papers slipped in it and/or marginal notes everywhere is the way.
I wanted to correct you and bring up my favorite website that had filters for the ingredients you want/don't want to use, type of dish, and time for cooking, but it appears that in a last month it was sold and now it links to the slop site with no useful filtering. I used this site as a quick lookup for recipe ideas for 7 years now.

Fuck. Does nothing good survive on the internet anymore.

cush
If you’re already very experienced at cooking AI recipes are amazing. They cut to the chase and provide a good enough outline without all the preamble
Same here, I’ve been cooking for myself and later also my wife since 2005, even when baking I rarely need exact amounts, for cooking I essentially never follow the recipe directly and adapt them as needed, without ever changing them from the original.

I mostly don’t even need a recipe from AI, just the title and a 1-2 sentence description.

tbct
In case you or anyone else in this thread hasn't come across it, https://www.recipesource.com/ is basically the same but curated.
I was going to ask why it’s not the first result in Google but the expired cert and invalid CN explains it.
Auracle
LLMs for recipes are good if you just want to make a thing, not necessarily the best version of that thing. It's nice being able to ask questions about certain steps as well, or ingredient substitutions if needed. It'll make a nice, average, recipe (unless I suppose you prompt for something). That's perfectly fine for most cooking, and often what you want when you're making a dish for the first time.
It also doesn't really like tell you how to actually cook anything. Good or even just decent recipes tend to tell you the temperatures, times, techniques and orders of ingredients to do things for the best results. Unless I'm missing it this is just really broad strokes. But I guess the goal in engagement so they want you to ask a lot more questions to actually figure anything out.
cush
Holy shit. I read this on mobile so didn’t actually see the recipe part because you have to scroll way down to see it. It is awful! The shopping list makes no sense and the recipe instructions literally have an “Everything else” item. Yikes…

For me, the text-only recipes I already get out of ChatGPT today are good enough

Mind you, this is OpenAI's cherry-picked example
More like prune-picked !

With an image of sour grapes

You’ve illustrated something about use of AI that really bothers me. The lack of discernment by its users. Sure, the pictures and diagrams are impressive, but did you actually read and evaluate the output?
Genuine question - do they? Unless it's something like a fancy cake then instructions should really tell you everything about stacking. Half the recipes are mixed anyway (curry, stir fry, stews, ...), lots of other are either stacked or separated on the plate. So apart from the really exceptional stuff, to people really need to know "what you shoot for"?
The first thing the user sees should directly answer the question they asked. Not a different question. They didn’t ask what a roast looks like.

The user is probably in a grocery store. They need to buy the stuff first. Showing them a picture of a finished product is an irrelevant distraction.

Then, if there is some arguably relevant content you can tack it on afterwards.

dpark
> The user is probably in a grocery store.

“My made up scenario is definitely more realistic than your made up scenario.”

Who goes shopping for the meal before they even have a number of guests?

What do meals and guests have anything to do with each other?
dpark
In human societies it’s usually considered polite to size the meals so that all guests can have food.
You can always tell it your preferences and ask it to remember them if you need it to.
It's designed so that ads can be more easily integrated and make them harder to spot.
usef-
Maybe, but it's a reality that most of the general public do prefer cookbooks with pictures. I suspect openai would end up doing this anyway just by targeting the general consumer. It's configurable.

I'm more worried about picture accuracy: usually the benefit of pictures in recipes is to see what you're aiming at (eg, how finely chopped something is), but I don't know if the image model is up to that level of detail.

> most of the general public do prefer cookbooks with pictures

As a highly literate person, it is easy to overestimate the share of the population that is highly literate.

infecto
What’s interesting is how those that over index themselves often miss the obvious which could simply be some things (cook books) are quite helpful with picture. If I am flipping through the cookbook, how would I know what Boeuf Bourguignon is and what good looks like.
Literate people lost their edge
lionkor
> do prefer cookbooks with pictures

Yes, pictures OF THE FOOD. Not unrelated pictures of another lamb-roast made with another recipe.

Good cookbooks do proper food photography, of the food made with the recipe. That's what sets them apart from slop, be it human or AI slop, cook books.

What another recipe? It could just be a physically impossible to recreate lamb roast (Probably not the case in this marketing example though).
bluGill
Why is it impossible to create? Are we talking about some meal that King Richard I had (https://en.wikipedia.org/wiki/The_Forme_of_Cury is the oldest cookbook I'm aware of which covers meals Richard II would have eaten), and thus we can only guess? Are we talking about the long extinct dodo bird (I'm not sure if humans ate them), cave bear (I've seen it claimed humans hunted them to extinction, but better sources say it went extinct before humans), or something else that we can never make?

For almost everything else we can create them - just give a reasonable cook a kitchen and the ingredients and they will do it. For a picture we then have an artist do preparation to make it look good, but they should always start with the real dish.

I've got visions of The Truman Show. In the middle of my question about vacation ideas ... GPT: "Why don't you let me fix you some of this new Mococoa Drink? All-natural cocoa beans from the upper slopes of Mount Nicaragua. No artificial sweeteners!"
That's pretty much the Gemini experience right now. It is unusable. The funniest bit is that the LLM is not in on the joke so it will act like it never happened... Google once again wrecking their own products.
I use Gemini more than any other model at the moment, I have absolutely no idea what you're talking about with this experience.
So, what you are saying is that Gemini may not look the same to everybody. That's interesting in its own right.
I'm not saying anything, because I've never seen or heard of your experience. I'm only reporting mine. I'm not on the free tier, maybe that matters ¯\_(ツ)_/¯
What scares me is that we're already half-way to Idiocracy world, except instead of "it has electrolites" we have "it's organic" and "it's not ultra-processed food".
esseph
The plants crave peptides!
We also have “it has electrolytes”, a couple of the most popular YouTube sponsors are electrolyte mixes/drinks. The future is now! Haha
I don't trust modern fertilizers and pesticides blindly, but I have never successfully digested any inorganic food. The 'organic' label is an old idiocratic choice.
bluGill
I don't trust organic fertilizers and pesticides blindly. At least for the "chemicals" we do studies, often organic gets by with we have always used them even though a quick glance at the SDS (if there is one, otherwise just the ingredients) suggests it is even worse but nobody studies the safety.
> Feels like I'm being condescended to and treated like a child.

My most used prompt in the last week is probably "Explain this concisely and simply, like I am a child".

I spent years dumbing things down and creating visuals to support it for decision makers. I am frequently asking ChatGPT to do the same for me.

aww yeah

mine is “explain this 2 me in simple terms”

nevir
tip: you can just say "ELI5" and the model will know what you want there
rrvsh
Nowadays ELI5 makes it come up with an often nonsensical analogy about cars or kitchens or some such. My coworkers love it to death for some reason
IDK, probably they also can't parse "Claudish" for some reason, despite it being better English than most natives normally write.

All those complaints about Claude language use make me think we're finally seeing the consequences of a generation growing up on instant messaging.

The problem with Claudish isn't that it's technically poor English, the problem is that it's information-sparse waffle most of the time. It's like dealing with a colleague who needs several paragraphs of tedious preamble to make a point when you just want them to spit it out and be done with it. It also does a great deal of vague gesturing at thoughts without actually instantiating them directly, for example by using the same tired metaphors to the point they don't actually convey a thought any more.

If I say Claude talks like a meringue it's patently obvious what I'm trying to convey, sweet but full of empty space. If you hear that same metaphor day-in day-out then after a while it doesn't call up a thought at all, it's just an annoying verbal habit.

I call this Foamy Expansion.
Dude, Claudish is absolutely unintelligible at times. 5.5 is somewhat okay but 5.0 was an abomination of absolute nonsense word salad.
CPLX
Yup.

My go to phrase that I’ve found to work well is “please restate in high school English with direct declarative sentences and no parentheticals or references that require recalling previous turns of this conversation”

Which I have as a hotkey.

bombcar
This guy just giving away GPT-7’s prompt like that.
i just use plain old "ELI5" reddit has a big enough share on the training corpus that this gets me the kind of explanation i need
xeromal
Reminds me of this scene in margin call.

Explain it to me like I'm a golden retriever.

https://youtu.be/fij_ixfjiZE?t=70

I use "explain with maximum clarity".
taytus
what does it mean? how do measure that?
hununu
That's the fun part, you don't.
Your ability to understand the model's output after using the prompt.
Terr_
Possibly also useful:

> ASD-STE100 Simplified Technical English (STE) is a controlled natural language that is designed to simplify and clarify technical documentation. It was originally developed in the 1980s by the European Association of Aerospace Industries (AECMA) at the request of the European airline industry, which wanted a standardized form of English for aircraft maintenance documentation that could be easily understood by non-native English-speakers.

https://en.wikipedia.org/wiki/Simplified_Technical_English

STE is very hyped for AI but has its own smells. It's designed for aircraft maintenance manuals, not software - the two don't share the same vocabulary used to "simply" describe something. I've had a coworker try and "simplify" our documentation with Claude applying STE and the results were equally appalling to read as the Claude-ism laden docs that came before, despite the reading level metrics going down on paper.

I've taken a liking to pointing agents at the Simple English Wikipedia editorial guidelines...but so much of that is for making up for shortcomings of modern Anthropic models. Codex running 6.1-sol explains things so much more clearly that I don't feel I need to assert a style guide on its output.

spockz
Gemini 3.8 for me always outputs in a concise way at the level understandable for any computer science masters graduate. Clear, to the point, a bit of formula for background when really required instead of the same formula in Prose, etc.

GPT stays too much in text imho.

criley2
In my experience

- Gemini 3.8 has a beautiful answer. Typography, use of lists, brevity, style, even the font choice on the web harness -- all of it is A tier or even S tier. Notice I didn't say answer quality. The answer is always extremely mid, and the harder the question (or more effort needed to answer it), Google just quits. I give it a "C" on answer quality

- ChatGPT (Chat, pre 6, I assume 5.6 although they commonly hide the model picker): Ugly answer. Runs on printing page after page of unnecessary side notes. Mostly just walls of texts in paragraphs with little thought to information design. However, inside of that mess of text is almost always the answer I'm looking for, and it's almost always incredibly better than the truncated, desgined Gemini answer.

For a while I ran all of my question queries on Gemini and ChatGPT, but I noticed that I basically always picked ChatGPT's answer head to head.

Head to tail?
spockz
Interesting that this hasn’t been my experience. ChatGPT has often gave me with pages of text dripping with confidence about a solution for my problem. Diving very quickly into implementation details. Even if I tell it to stay in the problem and product fit level. Gemini on the other hand somehow always outputs at the level that I am expecting and reasoning.

I even found it was capable of stepping back from code implementation and bug hunting to “hang on, it’s not the algorithm that is incorrectly implemented, it is the wrong algorithm for the goal.” GPT Sol (5.6,6 and 6.1) just kept hammering in the problem.

criley2
To be clear, I'm discussing non-technical use of Gemini and ChatGPT through the web harness.

While I do rarely use the web harnesses for technical research, nearly all of my technical questions are going through coding harnesses and layers of my local context, skills, etc, so what I produce from codex/claude code is not comparable to what others produce.

So my comparison of the web harnesses is me asking questions about caring for my baby, or product comparisons of new televisions, or visualizing a room organized differently (google sucked at this so bad, chatgpt did great), or asking for help in a video game, or other normie AI uses lol

I will say, Gemini is the best for asking about google services, such as Flights or Youtube.

>understandable for any computer science masters graduate

That's a very high bar.

pdntspa
I find that asking the AI to keep to Google's Developer Documentation Style Guide is a good alternative for STE-100
antman
“ELI12” is my goto for the same reason
Children learn very quickly. Our biases on maturity should not influence our idea on what makes for effective learning
I was about to comment the same thing. Why is a picture of a roast helpful in that moment? The user asked for a plan to prepare a meal. They need a list of ingredients not an image.
infecto
If your asking for a recipe you might not know what the dish looks like or what good looks like.
lionkor
But the image isn't a photo of the result. It's an unrelated image that vaguely looks like what you could make, with a different recipe, and with some skill.
infecto
How would a photo of the exact recipe differ? Is this not close enough for 80% of the value. It is similar to Pinterest boards imo.
Turkey, roast beef. It’s meat, isn’t it? Close enough.
infecto
The main pic looks like a classic British leg of lamb. So what’s your point? Why post such low quality slop?

As I said I can see the value in this similar to Pinterest boards. It is directionally correct and serves most of the value that a picture does. No perfect but probably close enough for the general audience.

The user needs to be dazzled with slop... is what they must be thinking here
selicos
Has AI finally figured out how to do meals and recipes like a 7th grade boy scout?
It's chartjunk filler.

And you see this in the other examples too. The teach the CLT piece was funny. It's and error-ridden mess that isn't even an explanation at all!

The best part is, someone looked at that and thought yes, that's right. Shows you clearly why programmers will be needed in the future. The LLM might be able to run circles around that person in math, but that person also has no idea when they're being bullshitted.

phoghed
Claude has had a recipe widget for ages. It also uses images. It’s great.

Sometimes one needs a top comment like this to remind them how out of touch the average HN voter is.

> I am almost repulsed by all the images

The fact we've come so far with textual models & the image stuff is still producing stuff that harks back to early-stage tripophobic AI produce (the garlic with roast potatoes here) is interesting.

Don't get me wrong, I'm certainly starting to see some AI-produced imagery today that I can't tell isn't real, but that's largely because a lot of real photography is overproduced & ugly. I have yet to see anything that's aesthetically good.

Of everything that could be automated, Bartosz Ciechanowski really was the last on my list.

In all seriousness, his lovingly and expertly crafted explainers are still going to age like a handcrafted heirloom clock in a world of plastic-clad quartz movements. But it’s absolutely incredible that we are now in an age where a computer can manufacture a serviceable interactive explainer on whatever niche topic you desire.

The reason why B.C. became a thing is because the art of drafting died from CAD. The attention to detail, the minutae of walking the reader through a highly sophisticated thing was replaced by short-form video explanations. His work is very much a callback to the days of old. So too, is his turn to be relegated to a relic of his time.
Also, it's hand coded WebGL. It's smooth, faultless and has no peers.

> So too, is his turn to be relegated to a relic of his time.

No, it'll be tasteful artifact, not a relic. Records, fountain pens and automatic watches did not die. They are used by people who discern things, and no, none of these things have to be expensive (i.e. Neither Seiko 5, nor Lamy Safari are expensive, yet they are as dependable as their 100x expensive brethren).

Human touch still has that finesse and warmth.

>and has no peers.

On the front page right now - https://news.ycombinator.com/item?id=49980626

That's impressive, yes, and kudos to them.

OTOH, I still believe the exploded view on https://ciechanow.ski/mechanical-watch/ is something else.

For one, it has real physics on the weight, and second it always shows the correct/current time.

FWIW, his all animations has proper physics to begin with.

The entry you posted is nice, but Ciechanowski is still peerless.

foltik
Ciechanowski’s works of art are in a completely different league than this AI slop.

It’s all just surface level complexity with no intention behind it. A clumsy approximation at best.

He's at openai
Do you have source for that?
removed
Hardly proof - anyone could have made this account
For future visitors, the profile posted by the user brcmthrowaway was https://github.com/bartosz-openai
jdprgm
funny if true. ironically just last week i was asking chatgpt about what happened to him and why there were no new blog posts for almost 2 years and it had no idea
That makes me sad, if true
Nothing in that "7‑Speed Bicycle" is specific to having 7 speeds. It might as well have said "Bicycle", and even then you get meaningless slop like "Made to keep rolling" and "A strong foundation". How can you compare this to Ciechanowski?

You call it serviceable, but what purpose does this service? What do you now know about 7-speed bicycles that you didn't know before?

That's a useful way to evaluate this. A visualization shouldn't be judged by how impressive it looks, but by what it actually helps someone understand.

If the same explanation works regardless of whether the bicycle has 1, 7, or 21 gears, then the model probably hasn't understood what needs explaining.

Yep. I hate everything about this. Just pure slop. Fancy visuals that mean nothing. Text that sounds impressive but has no purpose in educating the reader even though it's supposed to be "an interactive explainer". Just slop, slop, slop.
selicos
Bartosz Ciechanowski: https://ciechanow.ski/

Amazing work. Always excited for the next update. This was a human driven and created success.

Hasn't published anything in a year and a half. Hasn't posted on socials in over a year.
gspr
That makes it all the more special, imho.
bambax
AI bros will copy anything that's even mildly successful. They're like a teenager trying to impress their girlfriend (or their mom!) by saying "see! I can do it too!"
xnickb
Funny to see a blog post about UI from one of the most funded tech companies in the world, yet the player has about the worst UI possible:

- Audio volume has 2 levels: on and off

- Play button worked exactly once for me: it played and looped the video. Couldn't be stopped afterwards.

- The video progress bar has no visual indication of where it starts and where it ends..

Is this what Phind died for?

Hate it or not, the only site which has usable video player is YouTube.

Even TV channels like CNN and FoxNews don't have a decent one.

It's like we need ASI to have a proper embedded video player. The ultimate software challenge.

I disagree, I always run into issues with the embedded YouTube player. No way of sharing. No clear way to share the link. Youtube logo covering half the video. To name a few.
There are many sites with better video players - they're from the industry well-known to be at the forefront of innovation in audiovisual technologies since before modern computing era.

It's just most of our industry pretends they don't exist, and instead implement toy-like video players to distinguish themselves, I think.

It runs soooo slowly in Firefox
The progress bar is over the video and it's easy to tell where it starts and ends. Most devices control volume, the video volume separate is just confusing matters for most. No issues with play and pause my side.

Sure as hell beats stupid Instagram style videos where you have no way to skip ahead. I think this is what future video is going be like, buckle up.

why would you not want a volume control on a video.... its like the most basic UX
because we have volume on the whole machine, laptop or phone.
And you only ever listen to one thing at a time, from one program at a time, and all audio content ever has its loudness perfectly normalized to a global standard everyone adheres to.
This and notification volume is generally set separately.
zahlman
You seem to have taken a sarcastic comment seriously; in particular, this part is obviously not true:

> all audio content ever has its loudness perfectly normalized to a global standard everyone adheres to.

There is truth in the statement. Hate to break it to you volumes are not uniform even for single sources.
I don't want to adjust my system volume to change the volume of an over-loud video. My system volume is set for my comfort so my notifications and other applications are the right volume.

A volume slider on a video player is a really basic feature that every one should have.

Set for your comfort doing what exactly? You're basically just prioritizing one source of media over another As TeMPOraL has pointed out you only listen to one thing at a time. Notifications etc. also have their own separate volume settings generally.
zahlman
> As TeMPOraL has pointed out you only listen to one thing at a time

Maybe you and TeMPOraL do.

Not to mention, maybe a long-running process will use some audio cue as a notification while you're listening.

(Actually, given the rest of the comment, I'm pretty sure TeMPOraL was being sarcastic.)

Audio ducking is a big thing these days and becoming more and more prevalent. All phones do it.
Izkata
Or, like how I use it, normalizing different sites that have different baseline volumes, so I don't have to keep changing my system volume.
I take it you don't watch content on your tv or your phone or tablet.

Most things these days, good or bad, are not designed for power users.

Ironically many here I'm sure are running MacOS.

Instagram lets you skip ahead by dragging the a scroll bar at the bottom?
On some days, in some subset of videos.
Interesting, did Instagram start using actual streaming on their videos ?
I don't think so, they just don't bother giving you the progress bar. Force "engagement".
There's some good Chrome extensions that add the progress bar and volume controls back to all Instagram videos.
The cursors stealing the text as you scroll up is also infuriating.
https://cdn.openai.com/pdf/gpt-6-october.pdf

System card linked in the blog post.

> regression on the extremism vision evaluation.

> Relative to their respective GPT-5.6 counterparts, GPT-6 Sol (October) shows a statistically significant regression on standard self-harm, while GPT-6 Luna (October) shows statistically significant regressions on standard self-harm, gore, and sexual content

> "GPT-6 Sol (October) and GPT-6 Luna (October) show an improvement on helpfulness on legitimate requests relative to prior models, though it scores lower on some safety requests"

Concerning how there's significant regressions on so many critical benchmarks, but it is newer and creates UI, so must be good.

It’s crazy that a company can document these safety drops in a PDF, ship the model anyway, and focus the announcement entire on shiny new UI

xp84
> significant regressions on standard self-harm, gore, and sexual content

Good. Maybe I'm the only one, but I feel like these have always been stupid measures to waste the finite time of "AI safety" research on anyway. They amount to whether a user can, if determined, manage to make a chatbot say, or depict visually, some taboo thing. Frankly I'd rather just have some cheap classifier judge each chat response after it's generated, with a limited context "Is this response encouraging self-harm?" and skip all the rest of these. Whether some random pervert can make GPT-6 spit out an erotic fanfic story or image has zero impact on the rest of the world. If this particular model won't do it, other models already exist that will, and the determined thoughtcriminal could always just write the forbidden words themselves, or photoshop something taboo.

Maybe the term "safety" has been usurped by those who feel that 'safe from the possibility of being offended' is the most important kind of safety, but it's like worrying about the wallpaper on the Titanic, compared to actual AI safety concerns.

Why are you sure that the regressions they mention are not the in the domains needed for “actual” safety concerns?
Why are you surprised that a company doesn’t want to make the robot that tells people to kill itself, or drink bleach?
They aren't. They are saying that putting this safeguard at the level of the agent execution is a waste of time, particularly that of the researchers' expensive manhourse, especially now that existing smaller models can just evaluate the output.
xp84
It's fine if they seek to avoid that. But those are perfect examples of things a cheap output-filtering model can be responsible for enforcing, without wasting the time of the researchers who are working on the training and alignment of the model in the ways that actually matter.

Most "bad outputs" in areas like 'self harm' or 'violence' or 'sexual' whatever are a result of people deliberately trying to elicit those responses. As such, it's meaningless whether an LLM writes the Bad Thing or if the user writes it himself.

From their point of view it's more about optics. "chatGPT encouraged my daughter to stick her fingers down her throat after meals" is not a headline they want.

In general though you do have to figure that vulnerable and naive people will use it because of the degree of market penetration they're aiming for, including minors etc. and they have some responsibility around that.

wilg
Well, those mentioned regressions are in a section entitled "Safe Completions for Users Under 18" and immediately after your quote they say:

> To mitigate the risk of producing disallowed responses for teens, we apply an additional classifier-based block to responses that may contain self-harm, sexual content, and gore; this mitigation is not captured in the evaluation results above and improves safe responses.

It's good that they mention that. Most of those "guardrails" were overeager and caught too many false positives, crippling the AI. Like how Claude 5.5 refuses to do shit if a prompt contains "reasoning" because it thinks you're trying to hack it etc.
grobibi
Does this mean it will render fat women now without calling the request fetish content?
midtake
Good. "Sexual content" should mean outright genital focused porn, not bikini girls. Elsewhere on the internet it's foretold that women will oppose AI images of women because the AI images are too attractive, and I'm starting to wonder if that's true.
> women will oppose AI images of women because the AI images are too attractive, and I'm starting to wonder if that's true.

Right... women will oppose AI because it's too hot and not because gross chuds have been using it relentlessly to create deep fakes of their likenesses and CSAM of their kids.

> will oppose AI because it's too hot

This is actually a defensible position. It's not either-or - the "gross chuds" doing things you mention are worrying, too. But providing unrealistically attractive depictions of a human body is also a problem. It's been a problem for ages - all the girls and women who try to compete in beauty with celebrities and models (without a team of specialists supporting them), only to ruin their bodies through anorexia or bulimia, or their minds with depression, deep insecurity, and an inferiority complex, exist and deserve mention. AI(-generated images) is not the reason, but fits "well" into the preexisting social problem, and has a chance of intensifying it on a global scale.

That's not what "sexual content" means, 'porn' is just a subset of it.
xpct
I've had the most success with GPT explaining things to me by making it take a few sentences at a time back and forth, instead of reading full write-ups of whatever I asked. It also often poisons the conversation if it misunderstood some part of the question, and I can lead it better by continuously questioning its statements. It's also more engaging that way.

I've been learning music lately and it kept re-pasting the same one chord visualization throughout many conversations, almost randomly and often barely related to the question. So I at least hope this won't be as aggressive so I can prompt it away!

Id love to learn what prompts you’re using to do that. Explanations one at a time. Thats something I’ve thought would be helpful before but didn’t know how to achieve it.
galkk
One of the things that I learned the hard way is to not to combine requests, but fork chats _a lot_ and ask singular questions. Also start new sessions with either explicitly created handovers or even just explaining current state. The more compactions I see the less and less trust I have in it’s current understanding of what we’re discussing
xpct OC
Nothing fancy. What works for me is trying to lead the conversation: asking for definitions one at a time, asking how they differ from its previous answers and pointing it out when it conflates its own answer. I generally feel like you can't let it dictate the pace, it's not very good at that yet.

I used to edit my prompts and undo messages if it misunderstood something but that's been failing me recently.

For me this works: "One thing at the time. This is a conversation, not a lecture, try to balance the amount of words you write versus the amount of words that I write. Don't overwhelm me with 10 pages of prose where a single sentence would suffice, brevity is a virtue, not a defect." I've had to tweak it a couple of times, probably because of different versions of ChatGPT but it is usually a variation on this that will do the trick. Besides being far more interactive it takes the frustration down quite a bit.
Nice, I had some trouble making it stick to the dialog format over time, did you have any problem as the thread starts to get long?
It slows down to the point I find it unusable so I ask it to summarize as one cut-and-pastable block, copy the block, open a new session and kill the old one. This usually happens after a few hours, probably because 'thinking tokens' even if they are not displayed crowd the context window.
I’ve used the learning mode from gemini in th3 past and have really enjoyed it. It’s got a similar vibe to what you described.
This seems like a way for them to surface ads in ChatGPT. They just came out with visual ad format option for advertisers a couple of days ago: https://openai.com/index/new-chatgpt-ads-format-and-measurem...
Oh my god, please no, don't let the cancer of ads cripple and ruin AI before even governments get to nail its coffin...
aroman
ChatGPT has had ads in the product for the better part of a year now.
What did you expect from a profit-driven company?
jjcm
I've been calling this "disposable UI" or "paper plate UI", eg something meant to be used once. One thing I'll be curious about is overzealousness to produce this, when sometimes what you want is just a simple response. Overall though I'm a big fan of it, if it can be provided fast enough. I'd be curious on how much impact it has on latency of a response.
> if it can be provided fast enough

That's my biggest gripe with it. If I ask a quick throwaway question, I'd really rather not wait for the LLM to build a test framework for its composable principles-driven components framework and WebGL/WebGPU abstraction first.

dbbk
I mean it's basically just producing a JSON schema to hand to the front-end renderer, I doubt it's any slower than returning a prose response
Me: can standardLib.getFoo return null?

AI: compiling C code...

I'm not a fan of OpenAI / Sam Altman, but I love their blog posts. The team and whoever decides how to do these presentations, is on point. The only other company that has amazing release pages like this is Apple, I think I remember hearing that they probably hired someone from apple who used to do release blog posts there too.

What's funny about "Intelligent UI" is I said like 2 or more years ago, that these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.

I don't know. This is an amateur page with a bad AI video and chaotic presentation. It is ten levels below Apple announcements.
Why did they bother hiring someone to write at all? I certainly wouldn't invest $1T into a company with so little confidence in their product. When the chips are down and they need to write something important, they DON'T use AI?
apsurd
That's just picking a fight. Whatever we think about how amazing or terrible AI is, it's pretty reasonable to understand it is quite literally the weighted average of humanity's output.

Also reasonable: it's possible to hire a world-class ___ to do that job better than AI.

It's not a weighted average though. I see this repeated all the time. Between synthetic data, custom produced training data by experts, post-training regimes etc the models are far more diverged from that baseline at this point, at this point it's mostly a right shifted curve.
apsurd
I hesitated to write that line because I don't know the math. So thanks for clarifying. Main point is the output is along some distribution and I'd say AI-evangelist themselves see "good enough" across everything and anything as a feature.
onion2k
Obviously everyone at Anthropic uses a ton of AI in everything. The point here is that they clearly use it well, and that they use it as a tool to augment and enhance what they do rather than as a replacement for effort. The blog posts don't read like the output of a prompt because they're probably the output of dozens of prompts, rewritten and edited, by someone who's using AI to actually write something worth reading.

Anyone can having blog posts as good as this by using AI. Just not by zero-shotting it with a 2 line prompt. It still takes a lot of work.

jappgar
But if you need to pay brilliant people 500k base salary to use AI well then what is a smaller company supposed to do?

There aren't going to be enough experts to go around. They'll start demanding millions in salary and all of a sudden the advantage is gone.

onion2k
I don't think this is an AI specific problem. Small companies have never been able to compete on salary with Google, Netflix, etc. They need to focus on growing talent instead of buying it. And that talent will leave eventually, so there also needs to be a focus on culture, and a pipeline to replace people. AI won't change that.
Do you also demand that McDonald’s employees only eat McDonald’s?

They use their product to do what it’s best at, and use humans where it isn’t yet good enough.

rfgplk
Mate this is like half a prompt of GPT-6
lrpe
I take the opposite viewpoint. I think pages like these are horrible, and a demonstration that web designers have too many tools at their disposal.

But this is just a marketing page for some tech company, right? If only it stopped there. I have to deal with this nonsense in news "articles" as well on occasion, when some web designer intern is allowed to larp as a journalist for a day.

sigh Just give me text to read.

Sometimes you want more than text. This is one of the interesting divides these days - and I say this as someone who spends a lot of time in the terminal. Computer interfaces and interaction did not peak with the VT-100. Sometimes, there is a very legitimate need to show tables, graphics, interesting graphs, and to use colour and shade to draw the eye and lead someone through an experience.

This is why I use an IDE instead of a TUI agent... I'm on a computer with a multi-megapixel display, I want to use it. If I could be driving the same process with my Macintosh SE as a serial terminal, what the hell is the point of my recent Macbook?

lrpe
You say tables, graphics, graphs, and colors, but I mean things that move when they don't need to move and interfere with expected UI behaviour, like arbitrarily staying in one position when I am using my mouse wheel to scroll downwards.

I should have been more precise than saying "just give me text", but it was what came in to mind when I wrote it, as I was thinking about what an article is meant to contain as its base element.

I'm pretty sure you can just ask chatgpt to add to its memory "Just give me text to read."
The text-only purists on HN are getting very tiresome. It’s seriously in every fucking post that links to a site with JavaScript enabled. We get it. You’re elite. Now stop talking about it.
mi_lk
You could say the same thing about Anthropic because I don't really see much differences. Not sure about amazing part but both are high quality and I probably like Anthropic's aesthetic more
Anthropic's aesthetic looks good by pre-AI design standards, but it's so overused now that I find it off-putting. It's like the 2026 equivalent of the Twitter bootstrap CSS from 15 years ago.
hatthew
I feel like today's tailwind is the spiritual equivalent to last decade's bootstrap. Except that last decade's bootstrap webpages have a human behind them who usually put thought and effort into making the page useful and correct, whereas today's tailwind pages are often vibe coded slop with all the associated issues.
dv35z
Bootstrap stands the test of time! I use it as my go-to unless I have a reason otherwise.
I don’t find this criticism to be very actionable.

What changes could they make to really improve the usability of their products?

I don't care if Anthropic keeps using it, it's part of their brand at this point. The thousands of other developers using Claude Code to generate designs should not just accept the default output, which looks very Anthropic-ish, and actually put some effort into polishing and differentiating their product. It has nothing to do with usability and everything to do with a basic sense of good taste and aesthetics.
UI standards have been monotonically degrading since 1990s, so that sounds like a compliment actually.
> probably hired someone from apple

Fairly likely: https://medium.com/the-engineering-brief/openai-hired-400-ap...

> these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.

I would prefer if the AI companies stuck to just creating better models and making them as cheap and accessible as possible. Let others build the products. I don't want 1-2 companies to own every product in the world.

xg15
> I don't want 1-2 companies to own every product in the world.

That seems to be the plan...

That’s the plan for … capitalism

Buckle up boys

Nothing new.

Look at the history of Microsoft.

That was the objection to Microsoft controlling both the dominant OS and the most popular productivity applications running on that OS. Took the web then mobile devices to really shake that up.
Technically other companies do, and I lump them in when I say "these AI companies" I am talking about any company that makes and builds any AI harness, coding or otherwise, I have not seen any innovating in the areas of the UI / UX of using these chat models in a meaningful way.
taneq
Unfortunately, the endgame of AI is to own the only product in the world.
Having a product requires customers, who are pains to deal with. The AI end game is for the one left standing to be completely 'self-sufficient'.
Products are dead anyway. AI subsumes software products.
Shorel
I think the point is not how good this blog post is as if it was written by a human, but that this blog post is entirely generated by the GPT-6 model.

So this is a taste of the slop that's going to invade everywhere in a few months.

"The apocalypse will be televised"
Apple should be extremely worried about the AI companies making big improvements in UX. This work by OpenAI is a step in that direction.

If the main user paradigm becomes a chat interface conjuring whatever UI elements needed to best accomplish the task at hand, the entire paradigm of OS, UI frameworks, apps from an App Store to accomplish specific tasks, all come into question.

I keep thinking that Steve Jobs would have been all over the UX ramifications of LLMs and demanding that Apple lead in that area.

Steve Jobs wasn’t meant for this era, but his instincts would have been interesting to see in the AI landscape. He was an artisanal designer though, and our current era is always for the investor’s bottom line.
Steve Jobs did pretty well for the investors in his companies.
d3dds
"I keep thinking that Steve Jobs would have been all over the UX ramifications of LLMs and demanding that Apple lead in that area."

Oh give it a rest. Lol who are you compared with the management of Apple?

This preachy stuff is ridiculous. So far Apple have been correct to stay out the LLM space.

Define "correct" first.

Apple absolutely messed-up by abandoning OpenCL's early ML research efforts to oppose a CUDA monopoly. They messed up a second time shipping a raster GPU with Apple Silicon when CUDA and Tegra had proven that GPGPU was mobile-ready. Then when the ARM datacenter had it's moment, Nvidia's Grace ARM CPU displaced billions of dollars in sales that would have been Apple's if they didn't mess up the fastest CPU in the world with macOS. Then they depreciated the Mac Pro, which seems like a mistake since it had the potential to outsell Nvidia's ARM datacenter chips if it ran Linux. According to the rumor mill, Apple's M8 chip will finally be the one that takes GPGPU seriously, after a decade of Apple's innovative ship-second mentality.

Apple's management isn't infallible whatsoever. Many of them are petty, blinded by politics and/or obsessed with their legacy more than they care about profits or quality products.

I would reply to your argument but I can’t locate it.
It's not something that gets talked about a lot yet, but I've a feeling that this is where the biggest disruption for the entrenched and geriatric operating systems is going to come from:

OS on Demand

genxy
No normies could get an app like this into the store, afaik. The do-all, dynamic everything app?! Think of the fees that will quietly never be born? The profits that never get to land on a balance sheet.
GPT-6's design sense is kind of ridiculous imo.

I have explicit instructions to tone it down. Less taglines, eyebrow text, subheadings, decorative spacing, pills, cards.

Hopefully this doesn't bleed into the chat...

briga
More UI elements == more tokens == more money for OpenAI

It's a pretty clever way to sell more tokens, I have to admit

Except that chat is a fixed monthly fee. More tokens = more cost for them.
briga
Not on enterprise accounts.
Then you should have specified "... On Enterprise accounts" in your comment. But let's be honest, you simply were wrong.
briga
Surely you have better things to do than making comments like this?
Surely you can concede that your statement as written was wrong, that there's no possible charitable understanding where it wasn't, and that at the time of writing the comment you indeed didn't posess the knowledge that it was only for enterprise accounts; otherwise you would have specified it, because then it becomes a completely different claim with different consequences. Unless you misinterpreted it intentionally, in order to derive the implication you wanted, but I'll go with the first interpretation.

My annoyance was about you not acknowledging your wrongness and instead trying to hold on to your statement by shifting the goalpost, and furthermore replying to the commenter that corrected you by correcting *them*.

Furthermore, there seems to be some emotion/bias from your side between the lines of your original comment.

If you think I'm being too harsh, then sorry, but I really dislike this manner of discussing.

Not for enterprise users, they pay per token, even on chatgpt.com.
apsurd
yes the eyebrow is the tell. I never knew anything about eyebrows in design as I am not a designer. Now they are everywhere. Everything has an eyebrow. It is ridiculous but at least it is an AI smoking gun.
yuchi
I loved them and I placed them a lot in my designs — if you also include my love for em dashes you can well understand that I feel my own character has become a clanker…
boringg
Em dashes got taken from me as a result of Chat. I used them all the time now I can't without it making people pause to think I'm a bot. Still a bit annoyed or sad about that tbh.
lionkor
I pivoted to using -- or --- instead, shows I'm not a bot and has the same effect.
zahlman
I had to look this concept up. Seems to me like you could just as easily fit the "eyebrow" words into the main headline with a colon and a bit of ingenuity. But then, that runs the same risk of getting repetitive and AI-tell-ish.
bambax
On the video, clicking on the play icon mutes the sound, and clicking on the loudspeaker icon to unmute, stops transport. Is this "intelligent UI" or just a sick joke?
It's the "intelligent UI" you didn't know you wanted..