Doom Captcha (2021)
vivirenremoto.github.io
vivirenremoto.github.io
The ASCII intermediate interpretation also seems unnecessary and very limiting. But perhaps that's to keep it near realtime, looks like 1 FPS?
And why run on a Mac? Why not a beefy PC with a GPU that can do the calculations faster?
Still, does seem like a fun challenge. Maybe with further tuning or training it can level up
That's weirdly inspiring! What other games can I make where the visuals are conceptually no more than a line of characters, but which can get macroexpanded into immersive graphics?
I think the only weird part about that is that certain letter-number pairs may be a single token with some other semantics in the model, and other letter-number pairs would be a pair of tokens. I think that could impact the performance of the model (but probably not by a huge amount).
Nothing you are saying is technically incorrect. But, optimal performance was not the goal. The goal was to see if this crazy stupid concept would actually work. And, it does!
And yeah I totally get not aiming for optimal performance. I think it would be interesting to see how a language model could perform with a format that is less visually catered though. Like, textually there is little association between columns, it's just a string of characters, and some of them happen to be newlime characters. A more densely packed encoding would play more into the logic and reasoning encoded into the model, rather than just trying to parse out ASCII art.
Reminded me of this one: http://random.irb.hr/signup.php
(for a site with a slightly higher profile this wouldn't be enough, but for a minor corner of the internet with no ill intent actually aimed at it that turned out to be enough to block the fuzzing "fill all the forms" spam)
I guess that means the captcha is doing its job, since running LLMs isn't very cheap or scalable. But any harder problem means you start filtering a significant chunk of human users. Based on the other replies to your comment, it seems that the questions at their current difficulty already stop a lot of human users, yet allow a determined attacker with the setup I described pass through easily.
Then I refreshed the page, and was hit with calculus involving trig functions.
-3 * 3 + (-3) = ?
Where's my calculator?
Edit: oh wait. It's "least". I really have no idea then :)
With it being so famously portable, I was expecting this to actually run Doom in the browser and complete a simple map.
This makes sense when I try to indoctrinate my teenager who grew up on Halo and Call of Duty. But I began noticing this hangup in the late 90s with friends my own age.
Easily achievable[0], thoroughly obnoxious[1]. Just like all captchas.
[0] God help you if you're on a touchscreen. [1] For most people. Especially after the novelty wears off.
I'm sure you could still do it, but personally I try to respect copyright strictly for any projects I'm going to share. It just feels annoying to have copyright nonsense hanging over me otherwise.
It's free to play, sure, but is it free to use the assets for whatever you feel like and redistribute on your website? At a guess: no.
This surprised me because I thought that id's original shareware releases actually had more permissive licenses than that. Maybe the original Commander Keen did.
I guess maybe id/ZeniMax/Microsoft could theoretically sue you. But in practice the shareware assets are used completely freely without issue all over the internet.
Or maybe a remote desktop into an OS with a sandboxed browser that runs a Windows VM that ...
Alien doing pull ups? Fine. 8 year old girl holding a Quantum Physics book in a dark alley? That's sus...
I'm not expecting it to last longer, but there really should be some decent fishing bots at this point.
DOOM Captcha - https://news.ycombinator.com/item?id=27264988 - May 2021 (173 comments)
But AFAICT there is basically 0 money in browser games now, which is why only romantics and masochists still work on them.
Still cool and unique though
EDIT: Not broken, just not obvious one must click the sound options to start. Still just a mouse gallery mini-game. Doubtful you'd even need AI to solve it.
https://gist.github.com/enlyth/a177e4102b0da37a73587e15dbd68...
This could be further optimized to not scan the whole screen, and faking some human like mouse movements shouldn't be that hard too
1. There's no lighting, so the enemies have specific, fixed pixel colours that don't appear in any of the backgrounds. Scan and target these.
2. Enemies appear in a specific zone in the canvas. Makes scan faster, combines with below.
If there's expected ambiguity one can a. detect a few interesting background properties by looking at pixels where enemies never appear (e.g corners), and/or b. use a couple of other pixels relative to the candidate match (maybe neighbours, maybe not, could just as well be 20px down, 10 left) to discriminate.
Side story: one day my team was tasked with doing textual document content recognition for some biz. Everyone was like "oh it's going to be $$$ to pull out CV+OCR and have the OCR learn the specific font".
Turns out the document in question was:
- an extremely standardised gov format
- produced only by gov administration
- of a known fixed, overall size with clear identifiable boundaries
- printing known, standardised list of fields at fixed position
- with a known, standard font specifically made for quick automatic recognition
- containing only /[A-Za-z0-9]/ chars (plus a few I can't recall, but essentially dash, plus, slash...)
- on a known, standardised background
- the only variable is the quality of the scan and the size parameters
So I put a file upload form, piped the image through some reasonable imagemagick filter sequence to turn it into a no-background monochrome, look for corners/borders, resize+rotate, scan through the image til I hit a black pixel, then look at pixel-lit/unlit patterns (think 7 segment display in reverse).Cobbled the thing in a couple afternoons, with a quick, simple UI to have the user crop/rotate the doc (putting it mostly upright). It was stupidly fast to run and success rate was very high. Interestingly enough the failure mode was very good as it could reliably tell "ok I can't make any sense out of this" vs OCR which claimed success but outputted gibberish.
You can get surprisingly far with very little when you have known knowns.
But yeah, there is always a way to optimize. Even if making a clean room implementation (ie not looking at the source of that DOOM captcha) you can easily narrow down a recognition to a couple of 2x2 blocks and just pattern match them against a known background (ie not a monster).
DataDome talk about detection: https://youtu.be/xJGBfSGIsjw
Didn't know a simple interface with a sound switch and a game start button can be designed this badly.
Click to start:
[sound on] [sound off] [Start with sound] [Start without sound]1. an action – this is usually a verb in the imperative mood (e.g. Reply, Save, Add to basket)
2. a status – those omit the verb and only specify the new state of an object, which might be a lot of things, like a noun (Spam), an adjective (Favourite) or maybe an adverb (In progress, Later)
3. a navigation item – on the Web, this is better represented as a link, so let's not go into this here
I would argue that "with/without sound" is a clear example of a status here.
Also links are not buttons. There's nothing to get into here. It's straight up wrong from every perspective to think of a link as a button even from an accessibility standpoint.
[ CLICK TO START ]
[x] Allow sound
Keep it simpleI would also argue the MathDoku problem is different. That sounds like a mode confusion type issue, where the user expects a certain level of automation but it has been disabled by the system without adequate feedback.
Didn’t know a simple demo (with disclaimers) from someone who is clearly doing something novel could be commented on this badly.
A user's feedback is one of the best things that can ever happen to your program, the worst is to never ever get used by anyone, and the second worse is to have the users walk away with no idea why.
I certainly was confused and had a hard time starting it. If a significant amount of people can't even figure out how to start the game, the problem isn't inconsequential.
I tapped "click to start" on my phone a few times, saw nothing happened and assumed it didn't work on mobile and tapped back to come read the comments. I am neuroatypical, though, maybe I don't count as human.
Same reaction here.
Getting so pedantic about a minor point seems like it does more to stifle creativity and innovation and that it does to help.
As opposed to standard "click the traffic light" type captchas which are almost impossible for modern AI to break.
I think the doom captcha is probably more secure than standard captchas simply by virtue of its obscurity.
... and for humans, sometimes :)
"Standard" captchas sometimes also bring up major philosophical questions like "what is a bicycle?".