HNHacker News
TopNewBestAskShowJobs

xomateix

145 karma · joined February 17, 2015

submissionscomments
xomateix··on Six Degrees of Wikipedia
This made me remember a webpage that using IMDB would calculate the degrees of separation between Nicolas Cage and the input person (actress, director...) of your choice.

I did a quick search and couldn't find it. Does anybody remember it? I'm wondering if my memory or my search skills are failing or if it has simply disappeared.

xomateix··on Ask HN: Who is hiring? (August 2021)
Intent HQ | Scala Engineer | London, NYC, Barcelona, Lisbon & Remote| Full or part time

At Intent HQ we are working in a close relationship with our clients to help them have a better understanding of their data so they are able to provide better services to their own customers. We are ~80 people from all over the world (we speak 15 different languages!) based in London, Barcelona and NYC. As we are small, we love sharing ideas and really like to work along the principles of valuing 'individuals and interactions', 'customer collaboration', 'responding to change' and 'working software'.

Tech stack: Scala, Typelevel stack (cats-effect, htt4s, fs2...), Cassandra, Elasticsearch, PostgreSQL, Kafka, Docker, Nomad, Terraform, Consul, Vault, AWS, TypeScript, React, Redux

We have several positions open: https://intenthq.teamtailor.com/

Salary ranges depend on location.

If you want more information feel free to drop me an email: albert (at) intenthq.com

xomateix··on Is GitHub Copilot a blessing, or a curse?
According to GitHub support, they didn't exclude any repo based on the license: https://news.ycombinator.com/item?id=27769440
xomateix··on Pronunciations for hexadecimal numbers (1968)
In Spanish most people call it "tilde" as well, but you can also call it "Virgulilla" [1]. I always call it like this just because of how it sounds, love that word.

[1] https://es.wikipedia.org/wiki/Virgulilla

xomateix··on Ask HN: How to Teach Coding?
I am of the opinion that you don't teach anything, it's the person that learns something.

So, I'd go for what others have already mentioned. Learn about the person, how they are, how they learn better, what are their interests.

Find something they like and enjoy, so they are motivated in learning.

Be ready and available to answer questions and adapt to their rhythm and needs.

Prepare materials and different options so they can chose their path as they go.

Facilitate them changing their mind, going back and forth, making their own mistakes.

In my case, for example, I learn by doing, and pair programing with somebody helps me a lot, but other people might prefer having a theoretical background first and will want to read a book before diving into coding.

EDIT: adding paragraphs for clarity

xomateix··on Show HN: Anon – A Unix Command to Anonymise Data
Hey, one of the co-maintainers here. Thanks for your comments.

>> rows to randomly sample ... hash (using ... 32 bits) the column ... mod the result by the [constant] value

> This is not random. It deterministically selects the same very predictable fraction of rows.

Yep, you are right. We didn't intend the sampling function to be part of the anonymisation but just something we tend to use and we thought it would be useful to have it.

Its objective is to pick a portion of the input data. No more.

>> UK format postcode (eg. W1W 8BE) and just keeps the outcode (eg. W1W)

>> Given a date, just keep the year

> Partial postal codes and dates quantized to the year are still very revealing. Combined with other data (such as a hashed name), the partial postal code may allow a lot of people to be uniquely identified.

You are absolutely right. Depending on the use case and your data, having the outcode, the city or the year might be very revealing. In some other cases even having decades or centuries might be revealing.

We don't pretend that each function provided applies to all use cases. But in certain use cases partial postcodes or years can be good enough.

>> Hash (SHA1) the input

> Hashing does not provide anonymity.

We are very aware of that. That's why we offer the option to add a salt (that the user of the tool can make as long as possible and throw away after the anonymisation process).

>> range

> This is the only feature that could provide anonymity, if it is used correctly to group large numbers of individuals into the same bucket. This is probably more difficult that it first appears.

We usually work with sets of data that are tens of millions of users. Choosing the right ranges and, specially, analysing the data and making sure you anonymise the outliers (by choosing your bottom and top ranges carefully) it's crucial.

Again, this tool is a hammer. We expect a person that understands about wood and nails to analyse their problem and use it.

xomateix··on Show HN: Anon – A Unix Command to Anonymise Data
Thanks for the idea. We don't support anonymisation of IP addresses because it's not in any of our use cases yet. But I've already added an issue to address it.
xomateix··on Show HN: Anon – A Unix Command to Anonymise Data
Hey, one of the co-maintainers of the project here. And the one that decided to use json.

I agree with you. there are other options for configuration that are much better than json (yaml, toml).

Main reason for chosing json was simplicity. This was my first project in go and I didn't want to spend much time in it either. I found an example that was using json and I saw that I didn't need any external library to decode it. I thought that was good enough, at least for now.

Will probably look into using a library that supports yaml/toml for configuration in the future.

xomateix··on Show HN: Anon – A Unix Command to Anonymise Data
Thanks for the tip. We plan to add an examples folder. We'll add a preview too.
xomateix··on Show HN: Anon – A Unix Command to Anonymise Data
In addition to what Nathan has said, I'd add we needed a simple native command line tool that could be dropped into any server and easily work along other unix tools like cat, gzip, cut...
xomateix··on Show HN: Rambler – A simple and language-independent SQL schema migration tool
It looks interesting and I like the simplicity of the tool, looking forward for seeing more databases supported.

Besides needing the jvm installed, are there any differences between rambler and flyway command line [1] that make it a more suitable choice?

[1] https://flywaydb.org/getstarted/firststeps/commandline

xomateix··on Postgres Count Performance
We were doing some benchmarks recently on different databases. Most of the queries were heavy aggregations (>1billion of rows) and, to be honest, we were a bit disappointed with the new parallel query support in pg, we were expecting much better performance.

While doing the benchmarks, we could see that citus was always taking full advantage of all the cores in the cluster, while postgres parallelization was not.

Disclaimer: I'm not a db expert and I don't have any relationship with either citus or postgres.

xomateix··on Heroku is down in parts of Europe
Sorry for the link, but there is not an official announce or anything in their status page (https://status.heroku.com/) yet.
xomateix··on Re: Why Uber Engineering Switched from Postgres to MySQL
Was down for me, cached version: https://webcache.googleusercontent.com/search?q=cache:https%...
xomateix··on Theories on the etymology of 'strawberry'
In spanish and catalan it's called `piña` and `pinya` (exact same pronunciation as ñ is ny in catalan) because its shape is similar to the `Conifer cone` (called pinya as well) [1].

[1] https://es.wikipedia.org/wiki/Pi%C3%B1a

xomateix··on Introducing WhatsApp's Desktop App
Of course, he is not saying that, he is just explaining how whatsapp is actually storing the messages.
xomateix··on Redis 3.2.0 is out
Just realised that http://download.redis.io/ is not available under https (neither is http://redis.io).

It may be worth downloading it from github.

xomateix··on NGS: Next Generation Unix Shell
Relevant xkcd: https://xkcd.com/927/
xomateix··on Google.com partially dangerous
Just for the record, I quickly tried duckduckgo and bing (using !bang) and the first result I got from both the official chrome webpage.
xomateix··on SECRET DOM DO Not USE OR YOU WILL BE FIRED
Maybe because the answer is a whole thread that includes comments and other pieces of context that a simple copy&paste of just an explanation wouldn't have?
xomateix··on Vivaldi Browser 1.0
There is a private window option.
xomateix··on GitLab Pages
According to their documentation [1] it does work with private repos.

[1] http://doc.gitlab.com/ee/pages/README.html

xomateix··on Angola’s Wikipedia Pirates Are Exposing the Problems with Digital Colonialism
That's strange, as I got several emails from Amazon telling me about the upgrade.
xomateix··on DuckDuckGo grew more than 70% this year
You don't need to, you can activate the region button on the top right [0].

It's not the same as Google, but it's good enough in many cases.

[0] https://duck.co/help/settings/regions

xomateix··on File format wiki
A bit off topic, but I would like to know why did they choose mediawiki as a platform?

Under my limited experience working with wikipedia and wikidata I've seen it's not the best option to a) store structured data and b) edit the pages (markdown is, imho, much better for that).

xomateix··on Labella.js – placing labels on a timeline without overlap
The simple example link from the github page (http://twitter.github.io/labella.js/easy.html) is a 404.
xomateix··on Ask HN: Who is hiring? (November 2015)
http://intenthq.com | Barcelona | Full Time | ONSITE (mostly)

We are looking for Scala developers that will be:

- In the core of our business, being responsible for our most valued core IP

- Making sense of huge amounts of data

- Solving problems that don't have yet a solution

- Developing clean, robust and scalable code in Scala

And, most of all, we promise you won't ever get bored and will be having fun doing what we like the most, creating.

We offer flexibility and occasional remote work.

Contact me for more information: albert at intenthq dot com

xomateix··on How Akka Streams can be used to process the Wikidata dump in parallel
As you comment, Freebase is bigger than Wikidata. It is 22GB compressed (250GB uncompressed) while Wikidata is 5GB compressed (49GB uncompressed) [1].

Said that, I believe the process described in the blog post is not loading the whole Wikidata dump into memory and it would work the same to process Freebase or even larger data dumps with your laptop.

From the post: How Akka Streams can be used to process the Wikidata dump in parallel and using constant memory with just your laptop.

[1] https://developers.google.com/freebase/data http://dumps.wikimedia.org/other/wikidata/

xomateix··on 16-bit computer processor is being built by hand, transistor by transistor
See yesterday's relevant discussion: https://news.ycombinator.com/item?id=9755742
xomateix··on Why we are leaving Dropbox
I've been using copy.com for a while and I'm quite happy with it. The main reason I chose it (at the time) was 20GB for free accounts, fair storage in shared folders and pricing (I have the 250GB account).

(Disclaimer: this is a referral link that will give us 5gb https://copy.com?r=b2yUAQ ;-))

Page 1 of 2Next →