HNHacker News
TopNewBestAskShowJobs

rw

2,967 karma · joined January 27, 2008

https://github.com/rw
submissionscomments
rw··on CRDTs are the future
Operational Transformation and Conflict-Free Replicated Datatypes are very different from each other.

As the author explains, OT relies on some ordering of system events, and CRDTs don't. That means CRDTs need to be commutative (and probably associative), and OT doesn't.

So, OT is less scalable but more powerful, and CRDTs are more scalable but less powerful (in theory).

It's sort of like comparing Paxos/Raft to Bittorrent.

(I am not an expert on OT.)

rw··on Stegasuras: Neural Linguistic Steganography
Stegasuras is convincing work and the quality looks excellent.

I wrote a steganographic tool in this same spirit back in 2011, called Plainsight.

Back then, we didn't have deep learning, and the "Imagenet moment for NLP" had yet to arrive.

My Python code, with examples, is here: https://github.com/rw/plainsight

Unlike the OP, my Plainsight algorithm is 100% invertible by construction, and accepts binary input. (I verified the inversion process with "roundtrip fuzzing", a technique I still use today.)

Plainsight uses each bit of the input message to generate tokens. Bits are used to decide how to traverse a Huffman-style n-gram tree, weighted by frequency. This tree of n-grams is the model used in both the encoding and decoding steps. The drawbacks to my method are that the output 1) can be verbose and 2) does not convince a human that it's plausible, except for short messages.

Stegasuras has orders-of-magnitude better output, and seems to solve the problems I couldn't solve eight years ago. I would venture that their new result has as much to do with advances in language modeling, as it does with the particulars of their encoding and decoding algorithms.

I'll also note that I'm glad these researchers were able to use grant money to do this work. As a non-academic, I applied for an AI Grant to support me in upgrading Plainsight to use deep learning, but I was turned away at the time.

Finally, one of the ideas I picked up back then is that spam can be used to contain secret messages. Send enough gibberish to enough people, with your intended recipient included, and you'll look like a spammer--not a spy:

   $ wget https://spamassassin.apache.org/publiccorpus/20030228_spam.tar.bz2
   $ tar -jxvf 20030228_spam.tar.bz2
   $ cat spam/0* > spam-corpus.txt

   $ echo "The Magic Words are Squeamish Ossifrage" | plainsight -m encipher -f spam-corpus.txt > spam_ciphertext
   
   $ cat spam_ciphertext
   (8.11.6/8.11.6) 3 (Normal) Internet can send e-mails until to transfer 26 10 [127.0.0.1]
   also include address from the most logical, mail business for your Car have a many our
   portals ESMTP Thu, 29 1.0 this letter on internet, <a style=3D"color: 0px; text/plain;
   cellspacing=3D"0" how quoted-printable about receiving you would like width=3D"15%"
   width=3D"15%" border="0" width="511" Date: Tue, 27 Thu, 19 26 because
   zzzz@localhost.spamassassin.taint.org for
   
   $ cat spam_ciphertext | plainsight -m decipher -f spam-corpus.txt
   Adding models:
   Model: spam-corpus.txt added in 2.57s (context == 2)
   input is "<stdin>", output is "<stdout>"   
   deciphering: 100% | 543.84  B/s | Time: 0:00:00
   
   The Magic Words are Squeamish Ossifrage
rw··on TimescaleDB vs. InfluxDB: built differently for time-series data
The TimescaleDB benchmark code is a fork of code I wrote, as an independent consultant, for InfluxData in 2016 and 2017. The purpose of my project was to rigorously compare InfluxDB and InfluxDB Enterprise to Cassandra, Elasticsearch, MongoDB, and OpenTSDB. It's called influxdb-comparisons and is an actively-maintained project on Github at [0]. I am no longer affiliated with InfluxData, and these are my own opinions.

I designed and built the influxdb-comparisons benchmark suite to be easy to understand for customers. From a technical perspective, it is simulation-based, verifiable, fast, fair, and extensible. In particular, I created the "use-case approach" so that, no matter how technical our benchmark reports got, customers could say to themselves: "I understand this!". For example, in the devops use-case, we generate data and queries from a realistic simulation of telemetry collected from a server fleet. Doing it this way creates benchmarking stories that appeal to a wide variety of both technical and nontechnical customers.

This user-first design of a benchmarking suite was a novel innovation, and was a large factor in the success of the project.

Another aspect of the project is that we tried to do right by the competition. That means that we spoke with experts (sometimes, the creators of the databases themselves) on how to best achieve our goals. In particular, I worked hard to make the Cassandra, Elasticsearch MongoDB, and OpenTSDB benchmarks show their respective databases in the best light possible. Concretely, each database was configured in a way that is 1) featureful, like InfluxDB, 2) fast at writes, 3) fast at reads, and 4) efficient with disk space.

As an example of my diligence in implementing this benchmark suite for InfluxData, I included a mechanism by which the benchmark query results can be verified for correctness across competing databases, to within floating point tolerances. This is important because, when building adapters for drastically different databases, it is easy to introduce bugs that could give a false advantage to one side or the other (e.g. by accidentally throwing data away, or by executing queries that don't range over the whole dataset).

I don't see that TimescaleDB is using the verification functionality I created. I encourage TimescaleDB to run query verification, and write up their benchmarking methods in detail, like I did here: [1].

I think it's great that TimescaleDB is taking these ideas and extending them. At InfluxData, we made the code open-source so that others could build and learn from our work. In that tradition, I hope that the ongoing discussion about how to do excellent benchmarking of time-series databases keeps evolving.

[0] https://github.com/influxdata/influxdb-comparisons (Note that others maintain this project now.)

[1] https://rwinslow.com/rwinslow-benchmark-tech-paper-influxdb-...

rw··on I'm Scott Aaronson, quantum computing/computational complexity researcher. AMA
Hi Scott, thank you for writing your blog all these years. Your Busy Beaver essay ignited my passion for computer science, especially in algorithm analysis, logic, undecidability, and probability theory. I used to be someone who only thought in code; thanks to you, I now also think in math.
rw··on Show HN: Diamond – Full-stack web-framework in D
Why the hard dependency on MySQL?
rw··on China using big data to detain people before crime is committed
"Your scientists were so preoccupied with whether or not they could, that they didn't stop to think if they should."

- Jeff Goldblum as Dr. Ian Malcolm in Jurassic Park

rw··on Show HN: Using LDA to suggest GitHub repositories based on what you have starred
Good idea, the READMEs would be best of all.
rw··on Show HN: Using LDA to suggest GitHub repositories based on what you have starred
As I said, your approach is a clever way to use the GitHub API. I think you need to change the title and readme to indicate that this isn't an LDA index of GitHub descriptions. To ML practitioners, that's what you are implying with a title of "Show HN: Using LDA to suggest GitHub repositories based on what you have starred".
rw··on Show HN: Using LDA to suggest GitHub repositories based on what you have starred
This only uses LDA on your starred repository descriptions, to find topic terms that describe your starred repositories. These topic terms are then used to query the GitHub search API to find matching repositories. The results are then sorted by star count.

That is a clever way to make use of a search API like GitHub's. The principled way to do this, though, is to run LDA over all descriptions on GitHub, then use that similarity index to find similar repositories. You could run LDA over code, too.

I'll note that there is a cold start problem with this implementation: using LDA on such a small set of short documents will often lead to uninformative topics with words that are too-specific. You need a big corpus to capture e.g. synonym relationships.

rw··on Gophersat: A SAT solver written in Go
No, polynomial time. For reference, see these Wikipedia pages:

https://en.wikipedia.org/wiki/Polynomial-time_reduction

https://en.wikipedia.org/wiki/Karp%27s_21_NP-complete_proble...

rw··on Zuckerberg's trust problem
Why has this article been removed from the top 250 news results? It was #1 for a few minutes, then #5, and now it's gone. We've successfully discussed much more risqué topics here on HN...

Why did the comment by `TAForObvReasons calling out this apparent censorship get deleted?

rw··on Thoughts on OpenAI, reinforcement learning, and killer robots
No, it's called insufficient feature engineering. Data leakage is when your test data contaminates your training data.
rw··on The Future of Go Summit – Ke Jie vs. AlphaGo
A) How would you characterize the differences and similarities between AlphaGo and the best human players?

B) How has human play style changed since AlphaGo's introduction?

C) What is the answer to the question you most want to be asked?

rw··on Open sourcing Sonnet – a new library for constructing neural networks
TensorFlow is a dataflow computation system. Keras is for building neural networks. Each exists at a different level of abstraction.
rw··on Cracking Minesweeper with Z3 SMT Solver
How does this contrast with Answer Set Programming (using e.g. clasp)?
rw··on Introducing Keybase Chat
How did you find these changes?
rw··on The Axiom of Choice Is Wrong (2007)
You could have answered all of your questions with "finitely many", because, after all, we can each only perform a finite number of actions in the world.

In general, the infinite hierarchy of infinite sets "exists" because we can define it.

rw··on Skip Lists Done Right
Interesting. Whether that approach is efficient depends entirely on the workload. That complicates the analysis. (And one of the major benefits of skiplists is that the analysis is supposed to be simple.)
rw··on Skip Lists Done Right
I've read this page before, and again today, and I still don't understand how these unrolled lists are supposed to work in practice.

Based on the author's example at https://i.imgur.com/FYpPQPh.png, how do you take an unrolled skiplist that has a bottom row like this:

    [1,2,3] -> [4,5,_] -> [7,8,9]
And insert 2.5? An inevitable tree restructuring would have to occur, which vastly complicates the insertion logic.
rw··on Quick, Draw
This is pulling from a too-small data set.

I was able to get correct detections with too little data: In no world is a circle with a bar coming out of it a "dumbbell".

rw··on Udacity open-sources additional driving data
The world is bigger than web companies.

For contrast, one experiment at CERN produces 40TB/sec of sensor data, before downsampling and filtering: https://en.wikipedia.org/wiki/Compact_Muon_Solenoid#Collecti...

rw··on Carbon nanotube transistors outperform silicon
> biggest energy source of our galaxy

You meant the Sun, of course. However, your phrasing sent me on a quest to answer the galactic question. From the following article, I learned that the "supermassive black hole at the centre of the Milky Way" emits "cosmic radiation at petaelectronvolt energies", which fits the bill for largest energy source of our galaxy.

https://www.mpg.de/10390310/acceleration-petaelectronvolt-pr...

rw··on Ask HN: How to get out of Tech and still make a decent living?
What part of tech are you in? Maybe you're burnt out on that niche's particular subculture, not tech in general. For example:

Tired of 10-person startups? Move to a bigger company.

Tired by the constant churn in front-end libraries? Move down the stack and work on server software.

rw··on Stanza: A New Optionally-Typed General Purpose Language from UC Berkeley
Under the documentation section for "Ambiguous Methods"[0], why does the example compile? I would expect it to barf at compile time, not runtime.

[0] http://lbstanza.org/chapter4.html#anchor52

rw··on Programming on Parallel Machines: GPU, Multicore, Clusters and More
I confusingly used 'data' in two senses there... the second was is in the sense of 'code is data'.
rw··on Programming on Parallel Machines: GPU, Multicore, Clusters and More
Is your thesis online where we can look at it?
rw··on Programming on Parallel Machines: GPU, Multicore, Clusters and More
GPUs are best under SIMD conditions: single instruction, multiple data. You're talking about running `eval` thousands of times. Each unit of execution is going to have different data, because each process is executing different code (especially when you consider different branches of a conditional statement).

So, it wouldn't work that well :-)

rw··on The Unreasonable Reputation of Neural Networks
Fair comparison, but I didn't intend any monopoly on fear. All I'm saying is that the risk of GAI should be taken seriously, calibrated along with the other problems you mention. There's a lot of room between 'irrelevant' and 'the most important issue' :-)
rw··on The Unreasonable Reputation of Neural Networks
I agree that the rate of progress in AI is unpredictable, which means we are probably not right on the cusp of superhuman AI. But what if we actually are on the precipice? How could you tell? You seem to be taking a bet on the following statement:

"Before superhuman AI is developed, the techniques that make it possible will look dangerous."

That's a risky proposition. There's so much to lose in this situation. Elon Musk, Stephen Hawking, and many other people are taking a different, more risk-averse bet:

"Superhuman AI is extremely dangerous, so we need to be pessimistic about how good we are at predicting when it will happen."

By that logic, general AI is to be feared and worried about right now, because our predictive abilities are imperfect. The AI trend is towards more danger, not less: there's a slight chance of a cataclysmic event happening today, and as time goes on, the likelihood of it happening will increase (due to ongoing R&D in AI).

(As an aside, if you want to learn how hard the "AI Control Problem" is, I recommend the book "Superintelligence" by Nick Bostrum.)

rw··on Facebook uses FlatBuffers on one billion Android devices
Please read the comments on the post. We addressed the misunderstanding.
Page 1 of 27Next →