Show HN: Jevman – AI decision models play Pac-Man

open
65 points by felix089 · 17 comments
Visit opper.ai
Openai just launched their decisions endpoint, cloudflare launched clef the other week, and many more jev alternatives are out there.

We wanted to put the popular ones to the test and thought Pac-Man is a good benchmark for simple and fast decision making.

So we let jev 1.13, kev, clef, clef flash, GPT-6 Luna and Laya play Pac-Man against bot ghosts.

The low latency of these models allows for real time play. We had each model play 100 games, published a leader board and open-sourced the repo so anyone can run their own model and join the ranking. Link to repo: https://github.com/opper-ai/jevman-benchmark/blob/main/CONTR...

You can also join the game and play as Pac-Man yourself, and the ghosts are the models, either a mix of models or all jev, kev, clef etc. A game costs about 2 cent, all models are running via my startup opper, and we added free credits for everyone to try.

It's pretty fun to play and surprisingly difficult to beat jev's highscore. Any feedback is more than welcome!

17 comments

hung
Shouldn't it be running at every tick of the game rather than just the junctions? Pac-Man should be able to change direction at any point.
It may not matter if the ghosts' next move is predictable over that maximum distance.
That requires much more thinking (because of calculation) though doesn't it? Which decision models are not optimized for.
This is neat. At least I'm still better than the models at pac-man.

Also love the idea of a shared pool for users to try things out. I was considering more of a crowdfunded approach for one of my toy projects, something like... Giving it $10 in credits to begin and somehow allowing users to feed a buck in if they wanted.

yeah we had a lot of success with this also with another project, ai roundtable: https://opper.ai/ai-roundtable/history
felix089 OP
Great to hear, if it's about inference credits we actually built a thing called an AI wallet where users can just easily pay for their own inference on apps. if that's interesting lmk, id be happy to support this https://opper.ai/ai-wallet
nico
Very cool, are you also testing local, cpu-runnable/trainable models/classifiers?

I trained some to do some interesting things, including playing doom: https://github.com/nicobrenner/jeffy

I’ll try training one for this benchmark, seems like fun

Yes, try with toxic/toxichat and CFPB
felix089 OP
Not yet but the repo is open source and set up so that you can benchmark your own model and we can add it to the leaderboard, would be cool to see if Jeffy can play: https://github.com/opper-ai/jevman-benchmark/blob/main/CONTR...
Jev is useful to run locally overnight: it can classify the results while you sleep, ready for you to review in the morning.
Running cloud API inference for Pac-Man is peak modern software engineering. A 2 MB local ONNX model on a CPU core beats network round-trip jitter every time.
ozozozd
Super cool!

I wish the controls were a little easier on mobile.

felix089 OP
Thank you and yea agreed fixing it now, try again in an hour from now :)
it would be nice to have a regular LLM for reference
Paying 2 cents per game to avoid playing it.

What even is this reality.

haha thats one way to look at it :)
very cool :)