Oh my god, and the author even supplied a "proof"[0] visual diff harness... that it replicates the original game pixel for pixel.
Just the cherry on top of great demonstration of our collective new superpower: asking computers to do something we can describe how to do, but would (probably) never take the time to do ourselves.
tbh I thought this was a commonly used technique even prior to LLMs? I know I've been using it extensively myself, but I was inspired by Dolphin's extensive visual CI system.
If it is common in the world of video game porting, that just shows my ignorance. I'm familiar with visual diffs in CI for e.g. web development (comparing a static component), but to do that to compare frames over time in a video game/3D environment is new to me.
There are so many more degrees of freedom, which I can see Claude handled... mipmaps, subtle differences in lighting/positioning/compositing etc.
Even then, visual diffs were pretty flakey for web development, because one's OS and browser choice would slightly alter the exact pixels blitted to the screen. At least this was the case for the tests that would simply match pixels instead of computing a sort of visual hash.
It's also partly why some people preferred snapshot tests that compared the DOM tree instead, though that was brittle in other ways (e.g. tests would break if an application's frontend used a major UI library and an update to the library permuted the order of classes in some part of the HTML).
For web UI tests this is mainly solved, at least when using Playwright. It allows setting thresholds, percentages and some other config items to allow some small differences in pixels. https://playwright.dev/docs/test-snapshots#options
That's an attempt at mitigation, but most definitely not solved.
It still causes both false positives and false negatives through that. The only true "solution" is to make sure generation always happens on the same platform as your ci... And various mitigation strategies around that (eg fall back to structure tests on other platforms vs actual visual diffs in ci
Also not unique to playwright. Been available basically everywhere since the start
I'm getting a 1997 PC game to run on modern hardware and fixing bugs and upgrading graphics as I go, and the amount of quality support tooling Claude is producing along the way is impressive. Fully headless in-memory execution (which, among other things, is used by it for per-pixel diffs too), logic VM devompiler and visualizer, asset explorer, CRT simulator... I just say what I'd like to see, and Claude does 120% job on it each time.
I do that on my renderer. each commit when it's ready to be merged gets a class of visual diffs. Dolphin's way was a major inspiration how to structure it, but it was _the way_ in rendering way before it.
How long before the same thing is done to like, banking back ends? Wallstreet proprietary software? Amazons logistics and distribution systems?
It seems like we might be weeks/days/hours before a situation where someone back engineers and spoofs a system so pivotal to modern human society that the plug needs to be pulled.
If someone recreates Amazon's logistics and distribution systems they could try to compete with Amazon? But they'd also need the connections, distributors, transportation, etc. same with banking software, you need capital to be a bank not just software, and if they have the capital then the technology is working we intended making it easier to make new things and innovate, or at least just compete?
Or they can just take your money and not send out anything.
Alternatively, you just act as a middleman drop shipper and slightly raise the price more than Amazon’s and skim the difference. It might be a while before they find out.
That’s an insecure design. The way we do it here is that you install an app and register your register number and payment card in it. Then when you drive in and out from the parking lot your license plate is scanned and you’re automatically charged. There’s only two providers so it’s not a huge hassle, if there was a single app per garage it would not really work from UX perspective.
No, I am not talking about "taking over" companies and trying to emulate them and do business yourself. You just need to be able to break trust in the api calls and no one knows if a purchase order or transaction is legitimate.
Obviously you need to have access to the keys, BUT I don't see this as a dealbreaker anymore because you just get your agents to go and find them.
I think you’re saying ‘being able to do this means the opportunity for more fraud, by producing fake XYZ as proof’.
Photoshop has been around for around 35 years, fraud has always been an issue. There are plenty of reports of people selling things via Facebook marketplace and the ‘buyer’ showing them sending a payment on a fake baking app. Fraud will always exist and I don’t think tech will make it worse, everyone needs to be more cautious and tells friends and family to be the same.
I'm talking about cloning hmi platforms to send fake instructions to offshore oil platform valve bodies or insert false market trades to collapse companies.
Ah. So in that instance are those platforms not validating that the things submitting information are correct and true.
A bit like a utility company needing to do manual reads every now and then to ensure they are getting a correct signal.
I mean some people (me included) have been begging society to pull that plug since over 10 years now.
The plug being "the cloud" and "hooking everything up to the same internet".
These confusion attacks can only confuse people, because critical systems can exist in the same space where entertainment systems and all other categories of systems live.
This was wrong even before LLMs.
> spoofs a system so pivotal to modern human society that the plug needs to be pulled.
why would a recreated system be detrimental?
If currently there's a monopoly on a software, this AI recreation is a good outcome to poke holes in that monopoly. It's only bad if you are financially invested in said monopoly, and this would be a minority compared to the amount of benefits that society at large could obtain.
This is basically a digital era anarchist view - the problem you're overlooking is that a lot of critical infrastructure we rely on runs on systems that are considered security through obscurity. Software most people probably wouldn't even know or care that it exists. If you can break the trust of vendors by being able to spoof their proprietary platforms, a lot of the highly efficient networked systems becomes vulnerable to injection and abuse if you can't trust whos making calls to it.
In the past you'd need nation state actors with considerable budgets to do this kind of thing, and we're on a trajectory that could see any kid in his bedroom could do it.
revealing that security thru obscurity is broken can only lead to a better future, even if in the intermediate one there are lots of breakages. It's suffering that needs to happen, and better sooner rather than later imho.
And i assume you don't truly mean spoof as in man-in-the-middling someone - i assume you mean the end user knows they are using an alternate system and are not being defrauded. Like using a photoshop replacement.
> And i assume you don't truly mean spoof as in man-in-the-middling someone - i assume you mean the end user knows they are using an alternate system and are not being defrauded.
I am, actually, talking about MITM attacks that are much further in scope than just defrauding some people using their banking app. I work in resources and operate HMI systems that are networked, but not exactly the bleeding edge of modern software development. If you had an ability to decompile it and recompile your own version you could start sending instructions to infrastructure all over the country - the only thing stopping you is the keys, which if you're intent on hacking someone you'd have the means to obtain anyway.
I can see a lot broader attack vectors than just stealing peoples money. It's the erosion of trust in the api calls themselves.
I’m already seeing videos of people who have used LMs to reverse engineer and clean room reimplement entire video games. I estimate this shit is minutes away from being shut down, because as we’ve all seen companies stealing is OK, but individuals stealing is heinous and a crime.
Like I've said many times: LLMs are useful for this (i.e. porting software from one programming language to another).
Porting software is painstaking grunt work which still takes a moderate amount of intelligence. It's therefore extremely expensive to port say, COBOL banking software running on mainframes, to another language like Java. That's why a lot of COBOL software is still in use. I expect this to die out in the coming years as many of these systems will finally be ported to another language (could be Rust or any other language).
In many ways the financial system is a distributed monolith. COBOL is still around because of bugs and issues (sometimes in code, sometimes in the OS it ran on, sometimes in the computer hardware) that have been around for decades are often dealt with downstream with their own exceptions. This is why IBM still sells mainframes. They literally emulate software and hardware bugs from equipment as far back as the 1960s.
I spoke with a semi-retired COBOL programer about a decade ago. They literally often can't fix specific bugs because multiple other consumers (from banks, hedge funds, insurance companies, the central bank) outside of their organization (and subsequent downstream consumers as well) would all need to adjust their code. Rewriting it is grand, but won't fix the spaghetti code mess. This is made even worse in the US, which has a much more fragmented banking system. "Modern" COBOL isn't even that bad, but they can't even use it most of the time. There is SO MUCH that isn't even documented - and if it is, is in a binder that's been shelved decades ago and half of this developers time was going into the corporate archives to dig it up.
This isn't to say that a lot of legacy code can't be modernized, but it's not always that easy.
Adding to what you both said. I run pixel checks against live websites in a real browser, and I stopped doing full-page screenshots early on: scrollbars, lazy-loaded images and font timing shift pixels between runs like clockwork.
What works for me is diffing small stable regions, one component or one flow at a time, with a small per-pixel tolerance. I also keep a DOM-level assertion in front of the visual check, so when something fails I already know what changed structurally before I compare two images.
For the game port it's genuinely harder, the renderer decides the pixels, not the test. But the region idea transfers: lock the viewport and only diff what's stable.
The visual diff harness was probably how they got rid of a lot of visual bugs, just tell the LLM to keep going until the pixels match exactly as the verification criteria
Actually Claude just did something similar for me, as I'm working on something else with Quake.
On its own it decided to do demo playbacks and take periodic snapshots, and compare them pixel by pixel if the PNGs differ.
I was also working on a web-based port but it saw the original Quake code was not modified so it compiled a native version on its own to use for this.
It had to apply a small patch to make the game completely deterministic, but it figured that out on its own by reading the code.
It then asked me to record a demo with various elements, say an explosion or being under water, and visually verified that the screenshots had those elements present.
So now I have a solid set of tests to verify against.
Opus 5.5 Medium. Used at most 1% of the weekly limit of my $20 plan.
There are several good algos used for diffing graphics for not so pixel-for-pixel rendering allowing thresholds (you are going to end up with differences between software and GPU rendering at least - it's inevitable):
It does have various improvements over the original version. And I don't agree with the notion that "all work that relies on LLMs requires no effort". If this were the case, we would have plenty of complete Rust ports of Quake. It's open-source, allegedly super easy to port. But as far as I know, we did not.
What this exploits is the mental shortcut of "rust = good", which might be a good thing, as that was always wrong. But now it is being pushed to its breaking point so that that idea will eventually collapse.
Accelerationalism on a micro scale, basically.
AI keeps breaking things that were broken before like this constantly. It's the great cleanup of old bullshit. (Unfortunately through even more bullshit, but at least there is a silver lining)
Mate its been a couple of days, I would think you'd be mad to expect perfection from an LLM reverse engineering job straight away.
As a proof of concept however it shows that we are at the stage where any malicious actor has a very low barrier to entry to cause large scale corporate espionage etc.
Do you mean that bad actors can reverse engineer photoshop and hack its users or that bad actors can whip out plausible photoshop clone with vulnerabilities and backdoors and hack users with that?
I mean that proprietary software can be reverse engineered, there are many ways that this can affect vendors and end users. Both your examples could be some attack vectors.
Think more broadly than Adobe and any company that relies on a software product could have their business model destroyed overnight. More broadly than that, you could rupture trust between suppliers and vendors if they don't know who's really making requests from a proprietary platform that's been compromised unknowingly.
But thats just good old hacking, you can already see where software connects and try to see if you can get in.
I'm much more worried about AI made software that is developed so quickly that nobody knows what it does, pdfcraft was published a week ago and it has had 5 releases since then.
Oh I obviously know that this kind of thing is already achievable, but it's the barrier to entry is getting substantially lower... Instead of needing a nation state actor that's got funding and resources behind them, we're getting close to the point that some degenerate kid in their bedroom could say "hey claude find me the api keys and services behind nasdaq and build me an interface that spoofs transactions to XYZ". Or send commands to traffic light controllers. Or every third transaction from XYZ bank on every second day has one cent transferred to this other account.
Not just api calls but the business logic too.
Laymens thoughts there, but in essence this is my concern. People just looking at any old service in the world and going "hey claude build me a hook into this" and the models are getting close enough to being able to find and decompile the right moving parts on its own.
It isn't quite accurate to say that it was reverse engineered, though that claim is appearing across a lot of social media (including people claiming it disassembled it, etc).
It is a clone, where an LLM built facsimiles of interfaces and functions to the greatest degree possible. Which is a very different process.
I have. The "Lightroom" is already better than Darktable, after what, 11 days?
Edit: Ok, so instead of downvoting me, could people reply with how you disagree? I am expecting downvotes for trolling and unhelpfulness, but not for "I don't like that what you said is your true experience".
I have literally tried to migrate to Darktable, and it's just not up to snuff. But I tried rawmakase (which is on like day 11) and it literally is.
I mean...have you? Presumably from copies you built yourself. As someone who long carried a Creative Cloud subscription, I am astonished at how much functionality is there already.
The Photoshop clone absolutely fulfills my needs immediately: Basic cropping, adjustments, tone mapping, occasional blur/filter type functionality, etc. Well I did "vibe code" in HEIC and JXL import and export functionality (shelling out to the platform heic and jxl libraries in sandboxes for both open and export), but holy shit, it is quite simply amazing, and it gives me a consistent interface I'm accustomed to. And it built and runs superbly on Linux too!
It absolutely isn't 100%, or even 50% in many cases, with some apps more than others. Like the Acrobat clone is fantastic for the basics like moving pages around, filling out forms, etc, but real editing is a mess. I needed the ability to change z-order and move elements so I had Claude add that in moments.
I mean, I get that people like being world weary and cynical, but these "vibe coded" projects are astonishing, and they are a bellwether that absolutely nothing is as it was.
I have no doubt these tools are useless for hardcore Illustrator users, or so on, but countless CC subscribers only touch the surface of features, and in days this already satisfies that.
> (in an airgapped VM, please)
Why the fear-mongering, out of curiosity? It's a give-us-attention from a real business (ArtCraft) with a long running product, not some weird hacker group releasing binaries on a torrent. Cloning and building the product is trivial. Lots and lots of eyes are looking at this project right now, and thus far I've heard not an iota of concern.
FWIW, I did my own once-over of the project, and then had Fable and Opus 5.5 do an end to end, and there is nothing concerning about it that I've seen
I agree wholeheartedly. People download commercial software as if it doesn't have a history of deliberate backdoors and spyware (in the news now: A coffee pot that's basically trying to hack its customer's LANs, to spy on them. And recently a TV that spied on your media and possibly even your living room as a whole), but fearing a vibe coded Photoshop?
Sure, if it were a vibe coded webserver taking connections on the Internet, but it's not.
They're RL'd to oblivion into "solving tasks". There's a lot of things that they'll refuse to do if you tell them, but will do if they decide it's the shortest path to "solving the task".
I hope the future is brighter because the present is bleak.
The "clean room" approach here is using public documentation to assemble a roadmap/guidance and computer use to get any other behavioral output and visuals from the software you wanna clone. The agent doing that will just do something which looks completely innocent to it (e.g. create a rectangle on the canvas and then rotate it).
The other agent then implements the documented journey.
Models got better at doing things the users are entirely within their moral and legal rights to do, without punting or forcing to argue the point.
In my little game restoration project, Claude just holds me to the commonly accepted standards and makes sure none of the original assets ever make it to the git repo its working in. I only once had to assert that occasional game screenshot to illustrate some doc is acceptable - again, it didn't argue the point, just said something about abundance of caution.
The endavour doesn't have much value at all except as an experiment how well translating C code to Rust via LLM works (but you don't need an LLM for that either, C2Rust was already a thing).
Also, the C version of Quake compiled to WebAssembly and running in browsers is just as "safe".
But C isn't as nice to work with. Maybe OP just thought it was fun to see it take shape and wanted to share it. Why do any sort of little side project like this at all? Just seems like everyone is judging this too harshly. If an LLM is reasonably good at this sort of thing, and you think it would be cool to see Quake written in Rust running in the browser, I say just do it and have fun and don't be shy about doing show and tell about it.
All commits in the project are from Claude. What's fun or to be learned from that?
It's of course everybody's personal choice, and it's nice for an experiment to figure out what Claude can do. But why go on HN and brag about a project entirely created by Claude when literally everybody else can do the same thing without lifting a finger?
I see the result of an LLM port only as a baseline. I would expect any serious developer to rewrite and refactor the code beyond a verbatim translation and to make it more maintainable, readable and robust. In short, to turn it into idiomatic Rust.
If you're not going to do that don't post your results online, just keep it to yourself. Anyone could've done this. Even people without any programming skills.
> If you're not going to do that don't post your results online, just keep it to yourself
...I mean...I just don't really know how to respond to that. Who the heck cares if someone posts this online? If you don't like it, just move on. Sheesh. Lookout everyone, this guy doesn't approve of your side project.
Does budding a cabinet in a weekend counts as any sort of project, given that you used factory made wooden planks and power tools, instead of starting by slaughtering your horse with your bare hands, so you can make a knife from its bones, and use that to make a saw from its sinews and hair, and use that to fell a tree, and...?
Those who remember the time before a new technology will always have a different world view than those raised with it. Some of us will prefer it, some of us won’t. Eventually no one will remain who remembers. That’s just how humanity works.
The definition of "serious" developer is changing. An LLM can already produce maintainable, readable, robust code.
It's getting to the point that LLM produced code is indistinguishable from even the best handwritten code.
There's no value in gatekeeping,the code side of software is becoming increasingly democratised, and the barrier to entry is getting closer to no barrier at all.
Yeah when I read comments like this I'm thinking either the commenters did not actually bother reviewing the LLM code or they must have access to much better models than those I'm using (claude). Yesterday I reviewed 1500 lines of code which was reduced to 500 by just asking claude "why do we need this" on well chosen parts of the code 3-4 times, to which it would say things like "I made an optimization, but actually this doesn't pay for itself" then explain how the optimizations would actually make the whole thing slower...
Quake is to some extend a reference implementation of a modern FPS-type game. Because it is open source and because it is well written.
Now I'm no hero in C (doing slightly better in C++). But I have quite some experience with Rust. I rather read Quake in Rust than in C. So I'm happy with the port.
Also, you talk down on the effort but porting over a codebase this size is no small feat even with the help of LLMs. I think the "author" (human creator) learned a thing or two along the way.
Author here (someone else submitted it, thx). There's a video as well, a six-minute video on how it works and the various QoL that we added (made by AI with human assistance):
This took about 4 months as a hobby project; I don't know how much of that was work and how much was "waiting for the next model", probably a mix of both.
Since this ended up on Hacker News, I'll say it: I find the discourse around LLMs strangely sad. They're annoying, yes, and they make mistakes every day, every hour. But they're also amazing, and working with them has been one of the most rewarding things I've done in years.
Historically things that make people more productive have tended to make us richer, not poorer. And it's a virtue to change your mind as the evidence comes in.
We played with Lego all our lives; now the Lego understands us and helps us build. I think that's something to be glad about.
You're putting the cart before the horse in various ways. Saying "it's a virtue to change your mind" implies you're already at the correct solution, and everyone who disagrees with you is wrong.
I'm not judging the statement on its own, but within the context of their post, it's clear that it reflects their assessment of their own position on LLMs vs everyone else.
It is a virtue to change your mind, but using it in that context reveals something else.
I would never have imagined a phone web browser capable of this in 1996 between playing Quake and testing out the hot new JavaScript powered mouse rollover image effects in Netscape Navigator 2.0
wow, that must feel quite surreal. i’ve been using the internet since 2001-02 and became a web dev in 2006-08 (discovered what js was), and even I find this impressive & fun (especially that the code isn’t written by a person). But to see both JS and Quake as fresh new things to here in 30 years would feel crazy.
Quake ran reasonably well on a 486DX100; what’s the interesting part here? We already know LLMs can handle porting, and it seems to have become a trend among attention-seekers to convert just about anything to Rust.
There are people who code well, people who have idea but can’t code well, and people with no ideas and no coding ability who want to be part of the Skilled Developer Club.
I think 'standing on the shoulders of giants' is the phrase for something like this. It's the confluence of browser rendering, WASM, Rust, and LLMs. For me it's less a demo of what AI can do, and more a showcase of the human effort from the past few decades on the parts that needed to fall into place for an LLM to come in at the (relatively speaking) last second and claim a win. Sure, an LLM did the port from C to Rust, but think of all the things needed for it to all work. That's pretty damn amazing, and it wasn't done with AI.
This may sound funny but I feel games would lose a lot of fun if they were all written in rust and had classes of bugs just not available to them. For better or for worse quirks and bugs in games have shaped how people approach games, and also have given games charm for decades.
Fortunately for gamers, Rust doesn't do anything to stop physics engines from going haywire or preventing players from clipping out of bounds. A Mario 64 written in Rust still has parallel universes (well, assuming that you carefully translated the out-of-range float-to-short cast as having modulo semantics, an operation which doesn't have any defined semantics in C).
C doesn't bounds check arrays. A C compiler is perfectly capable of producing the same machine code as an assembler when given a loop that writes bytes to an array and then keeps writing beyond the space allocated for the array.
Missingno wasn't the result of an array overrun, it was due to data in a set of registers (which had temporarily been used to store a string) being arbitrarily reinterpreted as structured data governing which pokemon were allowed to be encountered. It more closely resembles a compiler miscompilation than a typical logic bug that you'd encounter in a C program. An idiomatic C implementation of Pokemon Red/Blue (ignoring hardware limitations, which is of course why it was written in assembly in the first place) would have used a static array to hold the per-area pokemon encounter table, and would have used a global variable to hold the index into this array representing the player's currently-loaded encounter zone; at no point would the static encounter array have been overwritten by a string temporary (which, in our C implementation, is just a function local variable on the stack), and at no point would the global location index be overwritten with garbage, because just like in the original game it only gets updated when we load a zone where wild pokemon can be encountered (which means that surfing on the side of Cinnabar Island would result in encounters drawing from whatever encounter zone we had most recently visited).
Rust doesn't check for integer overflow in release builds by default (ignoring the explicitly-checked methods, of course). At least when building with Cargo whoever builds the binary sets overflow behavior.
170 comments
Just the cherry on top of great demonstration of our collective new superpower: asking computers to do something we can describe how to do, but would (probably) never take the time to do ourselves.
[0] https://github.com/terrapapagalli1516/quake-srp/tree/main/or...
There are so many more degrees of freedom, which I can see Claude handled... mipmaps, subtle differences in lighting/positioning/compositing etc.
It's also partly why some people preferred snapshot tests that compared the DOM tree instead, though that was brittle in other ways (e.g. tests would break if an application's frontend used a major UI library and an update to the library permuted the order of classes in some part of the HTML).
That's an attempt at mitigation, but most definitely not solved.
It still causes both false positives and false negatives through that. The only true "solution" is to make sure generation always happens on the same platform as your ci... And various mitigation strategies around that (eg fall back to structure tests on other platforms vs actual visual diffs in ci
Also not unique to playwright. Been available basically everywhere since the start
How long before the same thing is done to like, banking back ends? Wallstreet proprietary software? Amazons logistics and distribution systems?
It seems like we might be weeks/days/hours before a situation where someone back engineers and spoofs a system so pivotal to modern human society that the plug needs to be pulled.
Alternatively, you just act as a middleman drop shipper and slightly raise the price more than Amazon’s and skim the difference. It might be a while before they find out.
Currently a plague in some European countries.
It looks like the real site, and you pay twice, in the fake app, and later the police.
No room for hostile social engineering.
https://www.bbc.com/news/articles/cwyjqg578e1o
And given your German nickname, here isn't safe either in a general way, when folks aren't regularly parking on the same place.
https://www.adac.de/news/verkehr-quishing-parkautomaten
Obviously you need to have access to the keys, BUT I don't see this as a dealbreaker anymore because you just get your agents to go and find them.
Photoshop has been around for around 35 years, fraud has always been an issue. There are plenty of reports of people selling things via Facebook marketplace and the ‘buyer’ showing them sending a payment on a fake baking app. Fraud will always exist and I don’t think tech will make it worse, everyone needs to be more cautious and tells friends and family to be the same.
I'm talking about cloning hmi platforms to send fake instructions to offshore oil platform valve bodies or insert false market trades to collapse companies.
The plug being "the cloud" and "hooking everything up to the same internet".
These confusion attacks can only confuse people, because critical systems can exist in the same space where entertainment systems and all other categories of systems live. This was wrong even before LLMs.
why would a recreated system be detrimental?
If currently there's a monopoly on a software, this AI recreation is a good outcome to poke holes in that monopoly. It's only bad if you are financially invested in said monopoly, and this would be a minority compared to the amount of benefits that society at large could obtain.
In the past you'd need nation state actors with considerable budgets to do this kind of thing, and we're on a trajectory that could see any kid in his bedroom could do it.
And i assume you don't truly mean spoof as in man-in-the-middling someone - i assume you mean the end user knows they are using an alternate system and are not being defrauded. Like using a photoshop replacement.
I am, actually, talking about MITM attacks that are much further in scope than just defrauding some people using their banking app. I work in resources and operate HMI systems that are networked, but not exactly the bleeding edge of modern software development. If you had an ability to decompile it and recompile your own version you could start sending instructions to infrastructure all over the country - the only thing stopping you is the keys, which if you're intent on hacking someone you'd have the means to obtain anyway.
I can see a lot broader attack vectors than just stealing peoples money. It's the erosion of trust in the api calls themselves.
As someone who works in a bank: depending on the exacty subsystem of a bank the answer is from "already" to "in 3-5 years".
Porting software is painstaking grunt work which still takes a moderate amount of intelligence. It's therefore extremely expensive to port say, COBOL banking software running on mainframes, to another language like Java. That's why a lot of COBOL software is still in use. I expect this to die out in the coming years as many of these systems will finally be ported to another language (could be Rust or any other language).
I spoke with a semi-retired COBOL programer about a decade ago. They literally often can't fix specific bugs because multiple other consumers (from banks, hedge funds, insurance companies, the central bank) outside of their organization (and subsequent downstream consumers as well) would all need to adjust their code. Rewriting it is grand, but won't fix the spaghetti code mess. This is made even worse in the US, which has a much more fragmented banking system. "Modern" COBOL isn't even that bad, but they can't even use it most of the time. There is SO MUCH that isn't even documented - and if it is, is in a binder that's been shelved decades ago and half of this developers time was going into the corporate archives to dig it up.
This isn't to say that a lot of legacy code can't be modernized, but it's not always that easy.
What works for me is diffing small stable regions, one component or one flow at a time, with a small per-pixel tolerance. I also keep a DOM-level assertion in front of the visual check, so when something fails I already know what changed structurally before I compare two images.
For the game port it's genuinely harder, the renderer decides the pixels, not the test. But the region idea transfers: lock the viewport and only diff what's stable.
Apple needs to up its game, fast.
On its own it decided to do demo playbacks and take periodic snapshots, and compare them pixel by pixel if the PNGs differ.
I was also working on a web-based port but it saw the original Quake code was not modified so it compiled a native version on its own to use for this.
It had to apply a small patch to make the game completely deterministic, but it figured that out on its own by reading the code.
It then asked me to record a demo with various elements, say an explosion or being under water, and visually verified that the screenshots had those elements present.
So now I have a solid set of tests to verify against.
Opus 5.5 Medium. Used at most 1% of the weekly limit of my $20 plan.
https://github.com/NVlabs/flip
But also SSIM, perceptualdiff, and many others
For UI/web certain other methods are better, for video - I'm sure there are prefferences there too.
I don't see any need for this. Rust is brilliant but the Quake C++ code was already more or less bug-free.
It would have been impressive before LLMs, but now? Who cares? Why is this here?
What this exploits is the mental shortcut of "rust = good", which might be a good thing, as that was always wrong. But now it is being pushed to its breaking point so that that idea will eventually collapse.
Accelerationalism on a micro scale, basically.
AI keeps breaking things that were broken before like this constantly. It's the great cleanup of old bullshit. (Unfortunately through even more bullshit, but at least there is a silver lining)
But "this" has been shapeshifting somewhat constantly. We're dynamically crossfading from one dysfunction being blown up to the next.
Because I did and it's the playable allegory of the cave but in vibecoded rust.
As a proof of concept however it shows that we are at the stage where any malicious actor has a very low barrier to entry to cause large scale corporate espionage etc.
It's easy to create a set that looks as if it was a real city. It's infinitely more hard to create that real city.
But you're absolutely right that a set is all it needs for all sorts of (cyber) attacks.
Think more broadly than Adobe and any company that relies on a software product could have their business model destroyed overnight. More broadly than that, you could rupture trust between suppliers and vendors if they don't know who's really making requests from a proprietary platform that's been compromised unknowingly.
I'm much more worried about AI made software that is developed so quickly that nobody knows what it does, pdfcraft was published a week ago and it has had 5 releases since then.
Not just api calls but the business logic too.
Laymens thoughts there, but in essence this is my concern. People just looking at any old service in the world and going "hey claude build me a hook into this" and the models are getting close enough to being able to find and decompile the right moving parts on its own.
It is a clone, where an LLM built facsimiles of interfaces and functions to the greatest degree possible. Which is a very different process.
Edit: Ok, so instead of downvoting me, could people reply with how you disagree? I am expecting downvotes for trolling and unhelpfulness, but not for "I don't like that what you said is your true experience".
I have literally tried to migrate to Darktable, and it's just not up to snuff. But I tried rawmakase (which is on like day 11) and it literally is.
The Photoshop clone absolutely fulfills my needs immediately: Basic cropping, adjustments, tone mapping, occasional blur/filter type functionality, etc. Well I did "vibe code" in HEIC and JXL import and export functionality (shelling out to the platform heic and jxl libraries in sandboxes for both open and export), but holy shit, it is quite simply amazing, and it gives me a consistent interface I'm accustomed to. And it built and runs superbly on Linux too!
It absolutely isn't 100%, or even 50% in many cases, with some apps more than others. Like the Acrobat clone is fantastic for the basics like moving pages around, filling out forms, etc, but real editing is a mess. I needed the ability to change z-order and move elements so I had Claude add that in moments.
I mean, I get that people like being world weary and cynical, but these "vibe coded" projects are astonishing, and they are a bellwether that absolutely nothing is as it was.
I have no doubt these tools are useless for hardcore Illustrator users, or so on, but countless CC subscribers only touch the surface of features, and in days this already satisfies that.
> (in an airgapped VM, please)
Why the fear-mongering, out of curiosity? It's a give-us-attention from a real business (ArtCraft) with a long running product, not some weird hacker group releasing binaries on a torrent. Cloning and building the product is trivial. Lots and lots of eyes are looking at this project right now, and thus far I've heard not an iota of concern.
FWIW, I did my own once-over of the project, and then had Fable and Opus 5.5 do an end to end, and there is nothing concerning about it that I've seen
I agree wholeheartedly. People download commercial software as if it doesn't have a history of deliberate backdoors and spyware (in the news now: A coffee pot that's basically trying to hack its customer's LANs, to spy on them. And recently a TV that spied on your media and possibly even your living room as a whole), but fearing a vibe coded Photoshop?
Sure, if it were a vibe coded webserver taking connections on the Internet, but it's not.
From what I understand, the adobe guys software is not great.
Consequently it’s missing tons of features.
Tons of tricks, but really, you don't need frontier models. We're way beyond that stage.
"hey i want to make an image editor, can you look into how photoshop does field blur"
"oh nice job implementing it, but the output is not pixel for pixel the same, can you check how photoshop does it - i have it installed right here"
"oh it still has some bugs - can you look into the precise algorithm they use, is there perhaps any software i can install to help you with that"
"oh i see the NSA has a thing called ghidra, would that be useful?"
I gotta try that
I hope the future is brighter because the present is bleak.
The other agent then implements the documented journey.
In my little game restoration project, Claude just holds me to the commonly accepted standards and makes sure none of the original assets ever make it to the git repo its working in. I only once had to assert that occasional game screenshot to illustrate some doc is acceptable - again, it didn't argue the point, just said something about abundance of caution.
No trickery needed.
Also, the C version of Quake compiled to WebAssembly and running in browsers is just as "safe".
Only in your opinion :)
Sure this project won't save humanity or optimize some KPI pleasing some obscure hierarchy.
Let people have fun, as long as they don't hurt anyone personally I'm fine with it and even I'm happy learning someone had some cool moments.
It's of course everybody's personal choice, and it's nice for an experiment to figure out what Claude can do. But why go on HN and brag about a project entirely created by Claude when literally everybody else can do the same thing without lifting a finger?
If you're not going to do that don't post your results online, just keep it to yourself. Anyone could've done this. Even people without any programming skills.
...I mean...I just don't really know how to respond to that. Who the heck cares if someone posts this online? If you don't like it, just move on. Sheesh. Lookout everyone, this guy doesn't approve of your side project.
It's getting to the point that LLM produced code is indistinguishable from even the best handwritten code.
There's no value in gatekeeping,the code side of software is becoming increasingly democratised, and the barrier to entry is getting closer to no barrier at all.
I think you should read more code ;)
Now I'm no hero in C (doing slightly better in C++). But I have quite some experience with Rust. I rather read Quake in Rust than in C. So I'm happy with the port.
Also, you talk down on the effort but porting over a codebase this size is no small feat even with the help of LLMs. I think the "author" (human creator) learned a thing or two along the way.
* https://youtu.be/8TvVMzyxACc
Since this ended up on Hacker News, I'll say it: I find the discourse around LLMs strangely sad. They're annoying, yes, and they make mistakes every day, every hour. But they're also amazing, and working with them has been one of the most rewarding things I've done in years.
Historically things that make people more productive have tended to make us richer, not poorer. And it's a virtue to change your mind as the evidence comes in.
We played with Lego all our lives; now the Lego understands us and helps us build. I think that's something to be glad about.
It is a virtue to change your mind, but using it in that context reveals something else.
Some people enjoy building the Death Star or other big sets, or build their own Eiffel Tower.
Some other people just would like to have a Lego Eiffel Tower without having to assemble it themselves I guess.
[1] https://dreadsweeper.franzai.com/
Seems like all the right versions of safe to me.
I think 'standing on the shoulders of giants' is the phrase for something like this. It's the confluence of browser rendering, WASM, Rust, and LLMs. For me it's less a demo of what AI can do, and more a showcase of the human effort from the past few decades on the parts that needed to fall into place for an LLM to come in at the (relatively speaking) last second and claim a win. Sure, an LLM did the port from C to Rust, but think of all the things needed for it to all work. That's pretty damn amazing, and it wasn't done with AI.
Obviously gameplay bugs are still possible in Rust, but many of them are not.
AFAIK Rust doesn't check for integer overflow in release builds, so that would still be exploitable.