HNHacker News
TopNewBestAskShowJobs

secretasiandan

626 karma · joined April 16, 2008

submissionscomments
secretasiandan··on Show HN: Xorq – open-source Python-first Pandas-style pipelines
Yes, "we" are out of core to the extent that the engines used in the deferred expressions we execute are out-of-core (our "batteries-included" engine is a modified Datafusion).

We have previously demonstrated the capability of doing iterative batch training by way of our "batteries-included" engine. I'll try to post a reference later but need to run now due to family obligations.

secretasiandan··on Show HN: Xorq – open-source Python-first Pandas-style pipelines
1. I would argue there are no "real alternatives". The two most proximate alternatives in feature space are Ibis and Snowpark.

- Ibis because while it can target multiple engines (as we state in our docs, we are built on and heavily reliant on Ibis), it aims to be "single engine, single session" in its execution in that nothing is expected to persist beyond the current session and an Ibis expression can only have a single engine. We want to be multi-engine and have some artifacts durable across sessions (by way of caching)

- Snowpark because it is sort of "multi-engine" by way of external functions or python stages, but locked to Snowflake. In some sense, we want to be Anypark: Snowpark like functionality but centered on whatever engine of choice is desired and performant interop with any other engines.

2. We don't have anything I would hold out as benchmarks yet. We don't aim to be "best in class" / the "fastest engine", we aim to be "in class" for as many operations as possible (we use the word performant). Our goal is to make it easy for an org to choose whichever engine(s) they feel most performant in when they consider the full space of {developer,computation} x {time,cost}. However, Hussain has demonstrated how having information from the "whole pipeline" available but execution deferred can allow for specialized optimization by way of predicate pushdowns (https://ibis-project.org/posts/udf-rewriting/)

Thanks for your interest and please feel free challenge any of the above or point us to anything you think we might have overlooked!

Best Dan

secretasiandan··on Why are so many coders still using Vim and Emacs?
A non trivial amount of on boarding is because "this is the most efficient information to ensure these learning agents have as they navigate their environment seeking rewards"
secretasiandan··on Eric Schmidt has applied to become a citizen of Cyprus
Honest question: what is the difference between taking advantage of the system and taking advantage of the people that comprise the system?
secretasiandan··on Shanghai stock exchange suspends Ant Group's A-share IPO
how do you determine the punitive capitalization ratio for non-capitalist regulation?
secretasiandan··on Uber CEO says its service will probably shut down temporarily in California
Have unicorns created an implicit threat of gambler's ruin with Uber taking the role of the house in this instance?

Let's call it "startup's ruin".

secretasiandan··on Briar Project
I don't think never worry about it again is quit correct

  What if someone registers with my old number? 
  If someone were to register with your old number on a new phone, then they will have an empty message history. Your contacts will also be made aware of a safety number change if they start messaging with the old number.

https://support.signal.org/hc/en-us/articles/360007062012-Ne...
secretasiandan··on Cracking down on research fraud
I think it is correct that on the surface it is credentials. I also think it is correct that the underlying property is social networks and a particular culture / mindset.

Ultimately, I think the "phonies" need to be addressed directly because I believe they are effectively a cabal.

I also think it requires the broader culture to take active steps to help make this happen: work harder to think about what bad behavior is and take action to avoid / penalize it.

FWIW, I have traveled in the finance / startup circles and kept looking for "better places" but have come to the conclusion that they are few and far between and the issue is the business culture and the broader culture that celebrates it.

secretasiandan··on Amazon scooped up data from its own sellers to launch competing products
Have you ever been tempted to tell people (journalists, government) about it? If not, are there any particular reasons?
secretasiandan··on Ask HN: Who is hiring? (April 2020)
Altana | Brooklyn, NY / Remote | Full Time

Altana is building a shared artificial intelligence platform to help governments, financial institutions, corporations, and logistics providers map and manage global flows of commerce, capital, data, and more. We have built the Altana Trade Knowledge Graph, the world’s most comprehensive representation of global commerce activity. This data asset covers more than 40% of cross-border transactions, corporate ownership registries in over 100 countries, the global movements of goods, illicit web activity, and more. Built on this foundation, our proprietary machine learning technologies and products are designed to help customers manage risk, automate otherwise labor-intensive investigations, and better manage cross-border flows.

Hiring for: Data Scientist, Machine Learning Engineer, Data Engineer See also: https://altana.ai/careers/

For Machine Learning Engineer or Data Engineer, email dan@altanatech.com For Data Scientist, email jobs@altanatech.com

secretasiandan··on White House Veterans Helped Gulf Monarchy Build Secret Surveillance Unit
This article claims it happens with lawyers too

  Mr. Pottinger later said that the scenario would have involved him representing a victim, settling a case and then representing the victim’s alleged abuser. He said it was within legal boundaries. (He also said he had meant to type “No client lawsuit is actually involved.”)
  Such legal arrangements are not unheard-of. Lawyers representing a former Fox News producer who had accused Bill O’Reilly of sexual harassment reached a settlement in which her lawyers agreed to work for Mr. O’Reilly after the dispute. But legal experts generally consider such setups to be unethical because they can create conflicts between the interests of the lawyers and their original clients.
https://www.nytimes.com/2019/11/30/business/david-boies-pott...
secretasiandan··on NUMA Siloing in the FreeBSD Network Stack [pdf]
I like it, I think its spreading, and I think its a good thing for the world.

"Top U.S. CEOs say companies should put social responsibility above profit" https://www.reuters.com/article/us-jp-morgan-business-roundt...

secretasiandan··on Ubuntu displays advertising in /etc/motd
You can watch HBO shows in amazon's player which works for chrome on ubuntu.
secretasiandan··on Differences between Tmux vs Screen (2015)
http://eclim.org/eclimd.html

> The most mature usage scenario that eclim provides, is the running of a headless eclipse server and communicating with that server inside of vim.

secretasiandan··on Zipf’s Law Arises Naturally When There Are Underlying, Unobserved Variables
From wikipedia: https://en.wikipedia.org/wiki/Zipf's_law

Thus the most frequent word will occur approximately twice as often as the second most frequent word, three times as often as the third most frequent word, etc.: the rank-frequency distribution is an inverse relation.

secretasiandan··on Program Your Finances: Command-Line Accounting
see also: http://furius.ca/beancount/
secretasiandan··on Ask HN: Who wants to be hired? (June 2014)
New York, Remote, Full Time | Contract | Part Time

Stack: Linux, AWS/EC2, VMware/VirtualBox, Python/Cython, C++, Bash, git, Jenkins

Resume: https://www.linkedin.com/pub/dan-lovell/4/655/922, https://github.com/dlovell

Contact: dlovell at alum dot mit dot edu

Overview: "data scientist" looking to develop quantitatively oriented systems requiring automation.

I am the primary developer of CrossCat (https://github.com/mit-probabilistic-computing-project/cross...), the backend of BayesDB (https://github.com/mit-probabilistic-computing-project/Bayes...).

I've developed infrastructure to

* programmatically create VMs for deployment, https://github.com/mit-probabilistic-computing-project/vm-in...

* programmatically deploy a Jenkins server to EC2 to run performance diagnostics for stochastic systems, https://github.com/mit-probabilistic-computing-project/jenki...

* numerical experiment infrastructure for performance evaluation https://github.com/mit-probabilistic-computing-project/exper...

secretasiandan··on Square Begins Offering Data Driven Cash Advances to Small Businesses
https://merchantfinancing.americanexpress.com/merchantfinanc...
secretasiandan··on Square Begins Offering Data Driven Cash Advances to Small Businesses
You don't need to be a bank to make a loan

http://www.sba.gov/community/blogs/community-blogs/small-bus...

secretasiandan··on America is no less socially mobile than it was a generation ago
From the paper: For example, if one defines mobility based on relative positions in the income distribution – e.g., a child’s prospects of rising from the bottom to the top quintile – then intergenerational mobility has remained unchanged in recent decades. If instead one defines mobility based on the probability that a child from a low-income family (e.g., the bottom 20%) reaches a fixed upper income threshold (e.g., $100,000), then mobility has increased because of the increase in inequality. However, the increase in inequality has also magnified the difference in expected incomes between children born to low (e.g., bottom-quintile) vs. high (top-quintile) income families. In this sense, mobility has fallen because a child’s income depends more heavily on her parents’ position in the income distribution today than in the past.
secretasiandan··on The Death Of Expertise
I think you're thinking of Isaiah Berlin's The Hedgehog and the Fox: https://en.wikipedia.org/wiki/The_Hedgehog_and_the_Fox

If you like that, you might also like the book Expert Political Judgement by Philip Tetlock which uses the Hedgehog and Fox metaphor to analyze people's political judgements. https://duckduckgo.com/?q=expert+political+judgement

secretasiandan··on Learning from Big Data: 40 million entities in context
What are the rows and columns in your hypothetical database?
secretasiandan··on Kevin Rose Will Join Google
I think you've entirely missed his point. Note that he says: "It doesn't matter if they understand how to code or not."

His main point? "Once they have that reputation as being awesome it will stick around no matter how badly they perform after their initial success"

And I totally agree. In fact, after a success the odds are more in your favor. So if you can't succeed again, it calls into question your 'true' level of talent.

secretasiandan··on Duck Duck Go Passed 1mm Searches Per Day
I don't think it matters how often you do !g, you're still going through ddg, which suggests ddg has a better relationship with you than google.

In facilitating your use of google for your specific searches, which you apparently can determine the likelihood of beforehand, ddg gets more of your searches for everything else. This is definitely a win for you and ddg, perhaps even for google if it increases your dependence on queries in general.

Also, I don't think ddg is trying to replace google. From your numbers, ddg is getting ~ 30% of your queries without having to implement deep search.

ddg doesn't log ip addresses/track you, so they could only probabilistically identify cases where people think ddg would do better but then fall back to google, but it will likely get harder and more computationally intense as they grow.

secretasiandan··on Stopping the Finance 'Brain Drain' Of The U.S. Economy
Perhaps that means companies buying the products that the financial sector is selling them should hire some smart people to work on their behalf.
secretasiandan··on Stopping the Finance 'Brain Drain' Of The U.S. Economy
If you work in an industry where this isn't true, but depends on the largess/spending of people who work in industries where this is true, are you still in a largely zero-sum waste of time job?

More simply, how many levels deep do you carry this analysis?

Also, can you guesstimate how much of the income generated in the US is not derived from zero-sum waste of time jobs/products when you carry the analysis at least 2 levels deep?

secretasiandan··on Facebook IPO To Raise $5B, Filing Wednesday, Morgan Stanley In Lead
Please explain "hedge with derivatives"

(Edit 1)

My issue with what you say : I believe options (the derivatives I assume you're talking about) don't start trading till a while after the stock starts trading.

Furthermore, even if they started trading when the stock starts trading, why do you think you have positive expected value to buy the stock and also buy puts? Why not just buy calls?

If either one of these is positive expected value, why do options market makers sell them at the price they do?

Additionally, your talk about the valuation seems trite : you're more likely to make money if you buy at a $50B valuation than if you buy at a $75B valuation is supposed to be informative?

(Edit 2)

Yes, I see your explanation, but you're (again) not really saying anything.

secretasiandan··on Facebook IPO To Raise $5B, Filing Wednesday, Morgan Stanley In Lead
One of the commentators on bloomberg (who asserted he always believed that the IPO size would be around $5B, on the lower end of the range) said he thought facebook credits would be a suprisingly large percent of total revenue and ad revenue would be lower than expected.

Reading between the lines: it may net to what everyone expects but the composition is important. It may suggest that Zynga is a good barometer for how facebook itself will fare.

secretasiandan··on Ask HN: Who's looking for employment?
Quantitatively oriented programmer (python,R,perl,java,excel/VBA,bash,SQL,matlab,linux/freebsd sysadmin, whatever)

Seeking remote work or on site in NYC (will move for the right opportunity)

Experience with quantitative trading models, parsing (EDGAR filings, web), competitive pricing analytics for online retailing, dynamical systems modelling

dlovell@alum.mit.edu

secretasiandan··on Cornell Said Chosen for NYC Engineering Campus
If you're asserting that Cornell has a poor engineering program, I think you need to do a little more research

"Among graduate engineering programs, Cornell was ranked 9th in the United States by U.S. News in 2008" https://en.wikipedia.org/wiki/Cornell_University

I don't think NYC needs a Stanford or MIT, they need someone who will treat the NYC campus as a primary focus. For that reason Columbia and Cornell should come first. You might argue that Columbia already has a presence but that doesn't mean they're NOT hurting for space. Furthermore, Cornell already has a NYC presence as well.

Page 1 of 5Next →