JSFuck – Write any JavaScript with 6 Characters: []()!+
jsfuck.com
jsfuck.com
One of the great points in the talk was JS obfuscation. Now, there are many techniques for doing this, but I really like this one as it just looks cool. Since you can translate most functional JS into ASCII, you simply encode every character into a binary coding using spaces for 0 and tabs for 1. Then you write a very simple converter + put an eval() around it, and viola, you are running arbitrary code. To a casual observer it would look like you have delivered a mostly empty file, while in fact you are delivering perfectly valid JS. Not unbreakable, but certainly fun.
Realistically, how big is your exploit? 30KB? So it'll be 30*7 = 210 KB (JS is actually 7 bit ASCII). That's plenty of code to do something malicious. Nothing is preventing you from minifying the code before converting it to this whitespace encoding.
I suppose you could make your encoding include other whitespace chars, like newlines, carriage returns, etc. Then you could use base 4 instead of base 2.
For a little bit of insight into why []+[] === '', for example, check this out: http://www.2ality.com/2012/01/object-plus-object.html
That code, run as-is in your browser, should alert 'hi'.
It's constructed with the intent of replacing the variables a-e and various named parameters with short sequences in the given character set that can be cast to strings. For legibility, I shortened many of the encodable sequences to strings (such as []["slice"])
The table holds 5^3 characters, which is 125; you could start the loop at 2 instead of 0 to get 2-127 as target characters, or use 6 input arguments. The 5 I chose were ![], [], +[], +{}, ~[] -- this should let you encode anything after overhead at a ratio of 12:1 with high repetition that gzip can take advantage of. This could be awful, I don't know; but maybe somebody will find this interesting.
[0] - http://patriciopalladino.com/blog/2012/08/09/non-alphanumeri...
It uses various techniques to get at e.g. type names, and then uses string indexing to access characters in the type name. It has enough characters like this to be able to call fromCharCode.
For instance, to get the string "a":
(![]+[])[+[[+!+[]]]]
Take the first part, `(![]+[])`. `![]` evaluates to `false`. Then `+[]` coerces false into a string, so the expression is `"false"`.The rest of the expression (more complicated) evaluates to `[[1]]`, which will grab the `"a"` from `"false"`. Now why there is the extra surrounding brackets, I'm not sure, because `[1]` would have worked as well.
(![]+[])[+!+[]] evaluates to "a" as well. It looks like it turns it into an array twice.
It goes +!+[] === 1 then [+!+[]] === [1] then +[[+!+[]]] === 1 then [+[[+!+[]]]] === [1]
Also confusing that "false"[[1]] works.
(Ok, tried it, it evaluates plenty of numbers on its own already which increases the character set. Not nice).
We start with [] and [[]].
convert [] to a boolean with !:
![] --> false
!![] --> true
convert [] to undefined by subscripting it: [][[]] --> undefined
convert [] [[]] and true to numbers by prefixing them with + and adding them with +: +[] --> 0
+[[]] --> NaN
+!![] --> +true --> 1
+!![]+(+!![]) --> +true+(+true) --> 1+(1) --> 2
etc.
convert any of these to strings by prefixing them with []+ []+[] --> ""
[]+![] --> []+false --> "false"
[]+[][[]] --> []+undefined --> "undefined"
[]+(+[]) --> []+0 --> "0"
get individual characters with the array subscript operator: ([]+![])[+[]] --> ([]+false)[0] --> "false"[0] --> "f"
([]+!![])[+!![]+(+!![])+(+!![])] --> ([]+true)[1+(1)+(1)] --> "true"[3] --> "e"
etc...
So now we can obtain a limited number of characters: "a" "d" "e" "f" "i" "l" "n" "r" "s" "t" "u" "N" "0" "1" "2" "3" "4" "5" "6" "7" "8" "9"
which we can combine into strings with the + operator: "f"+"i"+"l"+"t"+"e"+"r" --> "filter"
"1"+"e"+"1"+"0"+"0"+"0" --> "1e1000"
It's not much, but it's enough to obtain large numbers: +("1"+"2"+"3"+"4") --> +("1234") --> 1234
+("1"+"e"+"1"+"0"+"0"+"0") --> +("1e1000") --> Infinity
and, most importantly, to access a property of the array object: []["filter"] --> function filter() { [native code] }
and by converting these back to a string: []+[]["filter"] --> "function filter() { [native code] }"
[]+Infinity --> "Infinity"
we can expand our alphabet still further, and access some even more exciting properties: []["constructor"] --> function Array() { [native code] }
([]+[])["constructor"] --> function String() { [native code] }
(![])["constructor"] --> function Boolean() { [native code] }
(+[])["constructor"] --> function Number() { [native code] }
[]["filter"]["constructor"] --> function Function() { [native code] }
Almost there now.By converting these back to strings we can access even more characters, and by passing the strings to the Function() constructor we can can construct functions and evaluate them! In other words, we have "eval". Let's use it to access the window object:
[]["filter"]["constructor"]("return this") --> function anonymous() {return this}
[]["filter"]["constructor"]("return this")() --> window
So now we have access to the global context and eval. We don't quite have access to the full range of letters, but we have enough letters to call toString, and use it's base conversion ability to get the full lowercase alphabet: 10["toString"](36) --> "a"
11["toString"](36) --> "b"
...
25["toString"](36) --> "p"
...
35["toString"](36) --> "z"
And now we have "p", we can use escape and unescape to get most of the rest: unescape(escape(" ")[0]+4+0) --> "@"
So there you have it!The source code essentially runs this process backwards: it repeatedly uses regular expressions to convert the code back into "()[]!+" one step at a time.
It's more a showoff of a js idiosyncrasy - they found an ugly looking subset of characters that is Turing complete and wrote a translator to it.
If you're new to the concept of Turing tarpits, then this should blow your mind. On the other hand, this is a sufficiently advanced bit of CS thinking that no future employer should question the author's basic competency. I certainly have never written a translator.
You can get all the primitive values by taking advantage of unary plus, binary +, empty arrays, array dereferencing, function calls, and the standard strings returned by some basic expressions. The numbers are straightforward. The strings are dereferenced with numbers to get some individual letters. You can get methods by using array dereference on objects with strings. You get the rest of the letters with btoa and atob. Then you get eval, and you're off to the races!
I'm sure, but I think the encoder just goes token by token, eval'ing string conversions to get identifiers. You'll notice that for pretty simple expressions you get absurdly long strings out.
Well, okay, so it's still strictly speaking javascript but if we consider the functionally complete subset to be our 'target' language we end up in the same place, methinks :).
Is there a requirement that HN articles be ingenious? It's a novelty, that's all. And sometimes that's enough.
How exactly is this educational?
In order to get particular characters (for example: f), the script uses "false"[0], where "false" is derived from adding ![] + [], and 0 is derived from +[].
Putting all of that together ($ node):
> (![]+[])[+[]]
'f'
Here's some better examples of JS oddities: http://stackoverflow.com/q/9032856/1538708
> []["filter"]["constructor"]("alert(1)")()
The comparable IOCCC.org contest is a great place to learn C: sure the primary focus is making obfuscated code, and the audience marvels at the bizarreness thereof, but taking the time to pick a submission apart and figure out why it works (despite a host of reasons why you'd think it shouldn't) expands your understanding of the language.
It's a great way to improve your language skills when you've reached the initial "I know language X" confidence.
Tested a simple "for" loop: http://jsperf.com/jsfuck-perf-testa
Vanilla I got 225k ops/sec. JSFuck, 4.5 ops/sec.
So about 50000 times slower.
Which breaks down to: [!+[]+!+[]] === [2] +[[!+[]+!+[]]] === 2 [+[[!+[]+!+[]]]] === [2] //again
So I think there might be quite a bit of scope for compression even within the parameters it's built in.
Highly impressive, in any case.
http://utf-8.jp/public/jjencode.html http://utf-8.jp/public/aaencode.html
Then why does it still translate to 1128 characters? :/
Also, not cheap!
> []
[]
> +[]
0
> !+[]
true
> +!+[]
1
Lesson: + as a prepend coerces things into numbers > 123
123
> 123+[]
"123"
> ![]
false
> ![]+[]
"false"
Lesson: + as an infix operator between incompatible types coerces everything into strings > Function("alert(123)")
function anonymous() {
alert(123)
}
> Function("alert(123)")()
[alerts 123]
Lesson: the Function object can be called with a string to make a function that evaluates that string > [].constructor
function Array() { [native code] }
> ({}).constructor
function Object() { [native code] }
> (function(){}).constructor
function Function() { [native code] }
> (function(){}).constructor("alert(123)")()
[alerts 123]
Lesson: the constructor property returns the type of an object, and this gives you access to the Function object > [].filter.constructor
function Function() { [native code] }
> []["filter"]["constructor"]
function Function() { [native code] }
Lesson: an array's "filter" method is a function and has Function as its constructor. Well, duh. > [][[]]
undefined
> [][[]]+[]
"undefined"
> !![]+[]
"true"
> ![]+[]
"false"
> (!![]+[])[1]
"r"
> (!![]+[])[+!+[]]
"r"
Lesson: we now have the letters adefilnrstu, so we can construct []["filter"]. > [].filter+[]
"function filter() { [native code] }"
> ({})+[]
"[object Object]"
Lesson: we now have acdefijlnrstuv, so we can write []["filter"]["constructor"] and thus call our pseudo-eval on anything we can make as well. > Function("return console")()+[]
"[object Console]"
> Function("")+[]
"function anonymous() {
}"
> 0["constructor"]+[]
"function Number() { [native code] }"
> '0'["constructor"]+[]
"function String() { [native code] }"
> Function("return assert")()+[]
"function assert(condition, opt_message) {
'use strict';
if (!condition) {
var msg = 'Assertion failed';
if (opt_message)
msg = msg + ': ' + opt_message;
throw new Error(msg);
}
}"
Lesson: we now have all the letters we need to make ""["constructor"]["fromCharCode"]. > String.fromCharCode(74)
'J'
Lesson: Enjoy the rest of the alphabet. We can now construct arbitrary programs and eval them.Why aren't the six special characters defined as follows?
fuck()