HNHacker News
TopNewBestAskShowJobs

IvanVergiliev

86 karma · joined March 16, 2012

submissionscomments
IvanVergiliev··on Pg_hint_plan: Force PostgreSQL to execute query plans the way you want
Using the hint table has been pretty painful, in my experience. Two main difficulties I’ve seen: 1. The hint patterns are whitespace-sensitive. Accidentally put a tab instead of a couple of spaces, and you get a silent match failure. 2. There are different ways to encode query parameters - `?` for literals in the query text, `$n` for query parameters. Psql / pgcli don’t support parameterized queries so you can’t use them to iterate on your hints.

Still super useful when you have no other options though.

IvanVergiliev··on O(1) Build File
Just an FYI in case you haven’t read this: Recursive Make Considered Harmful [1]

[1] https://aegis.sourceforge.net/auug97.pdf

IvanVergiliev··on Anki-fy your life
You can use a “Custom Study Session” [1] for this. I tried it recently and it was fairly decent.

One possible way to get there on mobile: 1. Start reviewing a deck. 2. Click the gear button. 3. Choose “Custom study” and select one of the options. 4. If studying tagged cards, go tag some cards via the Browse view.

[1] https://docs.ankiweb.net/filtered-decks.html#custom-study

IvanVergiliev··on Accidentally exponential behavior in Spark
Definitely. One of the primary benefits we get out of Spark is the ability to decouple storage and compute, and to very easily scale out the compute.

Our main Spark workload is pretty spiky. We have low load during most of the day, and very high load at certain times - either system-wide, or because a large customer triggered an expensive operation. Using Spark as our distributed query engine allows us to quickly spin up new worker nodes and process the high load in a timely manner. We can then downsize the cluster again to keep our compute spend in check.

And just to provide some context on our data size, here's an article about how we use Citus at Heap - https://www.citusdata.com/customers/heap . We store close to a petabyte of data in our distributed Citus cluster. However, we've found Spark to be significantly better at queries with large result sets - our Connect product syncs a lot of data from our internal storage to customers' warehouses.

IvanVergiliev··on Accidentally exponential behavior in Spark
Post author here. Let me know if you have any questions!
IvanVergiliev··on Accidentally exponential behavior in Spark
Just responded to the parent comment as well - there's an additional mutable argument to the real `transform` method so it's unsafe to invoke it directly without first checking if the tree is convertible.
IvanVergiliev··on Accidentally exponential behavior in Spark
Good question, the simplified example doesn't make this clear.

The real implementation has a mutable `builder` argument used to gradually build the converted filter. If we perform the `transform().isDefined` call directly on the "main" builder, but the subtree turns out to not be convertible, we can mess up the state of the builder.

The second example from the post would look roughly like this:

  val transformedLeft = if (transform(tree.left, new Builder()).isDefined) {
    transform(tree.left, mainBuilder)
  } else None

Since the two `transform` invocations are different, we can't cache the result this way.

There's a more detailed explanation in the old comment to the method: https://github.com/apache/spark/pull/24068/files#diff-5de773... .

IvanVergiliev··on Old, Good Database Design
Exclude constraint on a GiST index?

https://stackoverflow.com/a/51247705

IvanVergiliev··on Ask HN: How do you learn complex, dense technical information?
Cool, I’ll take a read.

Just to clarify something in case it’s not obvious from my other comments. I’m not arguing for SRS as a replacement to all other forms of learning. You still need the extensive reading, problem solving, experimenting with a programming language.

However, I’ve found SRS to be a great addition to the methods above. For example, I haven’t found the methods above to give you long-term retention on their own. (And I’ve done a lot of problem solving.) Math is also a lot about building up the level of abstraction, and SRS can help with spacing the practice of lower-level concepts so you can more easily apply them to more complicated ones.

I’ve never been a fan of memorization in the past. However, I’ve found that:

1. Memorization (as in knowing foundational facts and being able to recall them efficiently) is actually pretty useful, as much as I didn’t want this to be true.

2. SRS can be used for thinking + deriving the answer to a card in addition to just memorization. It just gives you the right timing to do so.

IvanVergiliev··on Ask HN: How do you learn complex, dense technical information?
I’ve done some cloze deletions for math and related things, but I generally feel like having almost the whole proof in the prompt gives me too many cues. It often leaves me thinking that I indeed wouldn’t be able to come with the answer with fewer hints.

What I’ve tried the last couple of attempts is to “chunk” the proof (also terminology touched on in Barbara Oakley’s course) so that I end up with a question that’s something like “what’s the high-level idea / approach in the proof for X?”. That card would likely require an understanding of some underlying concepts or “chunks”, so I add questions for these too until I get to something that’s less abstract and easier to rederive.

I’m still not 100% confident if this will work well when these particular cards get into the 6-month range or so, and they start showing up at completely unrelated times. My main concern is that if I’ve forgotten some idea from “the middle”, it would be hard to reason about cards that build up on top of that.

IvanVergiliev··on Ask HN: How do you learn complex, dense technical information?
> it's a lot better to use those connections instead of drilling it in a decontextualized fragments via SRS.

Or, you can use SRS for "spaced repetition" of making these connections. That is, instead of treating it as rote memorization, use it for the timing effects. When I see a card about X1 which is part of a larger concept Y, I don't think "what was the exact answer to X1, which I remembered without any understanding and will just recite now?".

Instead, I often think "how do I come up with the answer to X1 right now? how does it connect to the larger concept Y?". Even better, if I've recently seen card X2 about the same concept, I might think "how does X1 relate to X2, which I just saw recently?". Sometimes, this actively helps you to make new connections. Of course, you need to explicitly make an effort to do so, yourself. If you practice pure recall only, that's what you'll get from SRS.

As another commenter mentioned, it's a false dichotomy.

IvanVergiliev··on Ask HN: How do you learn complex, dense technical information?
I'm a big fan of the idea of using Spaced Repetition for this. The idea is that it allows you to both:

- be able to keep your knowledge / understanding of an area around for the long term; but also

- be able to gradually build up your understanding by first committing the fundamentals to memory, and then using that to build up your level of abstraction and get to the more complex ideas and principles.

Michael Nielsen has written fairly extensive explanations of two slightly different approaches in:

- Using spaced repetition systems to see through a piece of mathematics [1]

- Augmenting Long-term Memory [2]

I've only used this particular approach for a handful of subjects so far - indeed, it seems to just take time to build a high-quality, long-term understanding of a thing. However, I've been pretty happy with the process so far.

[1] http://cognitivemedium.com/srs-mathematics [2] http://augmentingcognition.com/ltm.html

IvanVergiliev··on Focus has become more valuable than intelligence (2018)
I’ve found it extremely useful to keep a detailed log of my thoughts and ideas as I’m working on a problem that requires focus. It’s like a thread dump or memory dump of my thinking. Then, if I get interrupted for whatever reason, I can easily go back to the notes and “restore” from the thread dump.

This is a pretty good blog post I found on the subject: http://faq.sealedabstract.com/uninterruptible_programming_su... .

I’ve found various side benefits in addition to being able to focus in shorter time windows. For example:

- it’s useful for dealing with interruptions that are part of work too - e.g. if you’re helping teammates with different projects, or have to switch contexts for other reasons.

- it can be useful as an artifact of work. For example, you’ve spent a lot of time debugging a weird issue and you’re still not making progress, so you can use a second set of eyes. You can share your work notes with a coworker so they can immediately know what you’ve tried, what worked or didn’t, etc. In that context, I like to think of it as “offline pair programming”.

IvanVergiliev··on Distributed Filesystems for Deep Learning
> the only distributed filesystem with distributed metadata (apart from Google Collosus - the hidden jewel in their crown).

Isn't that the case for S3 as well? I haven't seen AWS mention it explicitly, but the consistency guarantees and the "unlimited scalability" claims seem to point in that direction.

IvanVergiliev··on Lisp in fewer than 200 lines of C
And throw an `inline` in there just to be more likely to end up with something macro-like.
IvanVergiliev··on Ask HN: What essay/blogpost do you keep going back to reread?
Yeah, I'm one of the top 1% readers on Pocket according to their stats, and my reading list has probably hundreds of articles. Didn't scale too well for me :)
IvanVergiliev··on Ask HN: What essay/blogpost do you keep going back to reread?
Meta-question: how do you keep track of these? Every now and then I come upon something and I think "Oh, I should definitely re-read this!". But then, eventually, I end up forgetting about it. Are these just things that often come up and re-read them because you know they're worth it, or do you actively keep track of them somehow?
IvanVergiliev··on Ask HN: What essay/blogpost do you keep going back to reread?
I've had this open in my browser for about two years now. I've read it maybe 1.5 times and I know I want to re-read it, but it's so long that I rarely actually get to it.

If you enjoy the subject, these two are good reads as well: - https://www.oreilly.com/ideas/the-world-beyond-batch-streami... - https://www.oreilly.com/ideas/the-world-beyond-batch-streami...

IvanVergiliev··on A Competitive Programmer's Handbook
Your second assumption is only valid for people who conflate the two disciplines. If you don't recognize that competitive programming and software engineers have completely different goals and requirements, then being trained in competitive programming may very well lead you to bad programming practices - e.g. single-letter names are usually fine for competitions, while at the same time usually horrible for engineering.

However, if you recognize the differences and use the skills you learn in the two as complementary, you can be better at both.

IvanVergiliev··on A Prettier JavaScript Formatter
Eventually we should just store the AST in source control. Prevents inconsistencies when style changes and/or 50k lines diffs to change the indentation level.
IvanVergiliev··on Scylla release: version 1.0
On the other hand, if the performance benefits prove to be true, you might be able to test Scylla with much less machines than you have for Cassandra.
IvanVergiliev··on Announcing Spark 1.6
Reduce can perform reductions on locally on each machine before shuffling the data. This decreases the memory as well as the network overhead. If you need all the elements for a given key - e.g. to display them to a user or save them to a DB, perhaps you should use groupBy. If you're going to perform some form of a reduce after that though, it's likely sub-optimal.