HNHacker News
TopNewBestAskShowJobs

ppsreejith

643 karma · joined February 13, 2014

https://ppsreejith.net
submissionscomments
ppsreejith··on Launch HN: MinusX (YC S24) – AI assistant for data tools like Jupyter/Metabase
Yep! MinusX uses Metabase APIs to pull relevant tables, schema, & dashboards to construct the context for your instruction.

> Anecdotally, my hardest problems w/ nl2sql are finding the right tables and adding the right filters.

Totally! especially in large orgs with thousands of tables. Using your existing dashboards and queries, gives useful context on picking the right tables for the query.

ppsreejith··on Launch HN: MinusX (YC S24) – AI assistant for data tools like Jupyter/Metabase
We're building an assistant that works across all your analytics apps. This means MinusX can use context from multiple apps to better fulfil your instructions. You can imagine a future version of MinusX reading data from a spreadsheet, putting it onto a Jupyter notebook / Metabase Table, and running further analysis.

When Metabase (or any other tool) builds an assistant, we aim to use it to further extend MinusX's capabilities!

ppsreejith··on Launch HN: MinusX (YC S24) – AI assistant for data tools like Jupyter/Metabase
Currently, we're using GPT-4o. We've tested it with Claude as well and plan to roll out support soon!
ppsreejith··on Launch HN: MinusX (YC S24) – AI assistant for data tools like Jupyter/Metabase
We've done a bunch of work to strip down the context and minimise the output tokens (which tends to be 100x as slow as input tokens). GPT-4o is pretty fast too :)
ppsreejith··on OpenAI: Cofounders Greg Brockman, John Schulman, along with others, to leave
Possibly related to Elon's new lawsuit? https://www.theverge.com/2024/8/5/24213557/elon-musk-openai-...

> The new lawsuit filed in federal court in Northern California on Monday says that Altman and Brockman “assiduously manipulated Musk into co-founding their spurious non-profit venture” by promising that OpenAI would be safer and more transparent than profit-driven alternatives. The suit claims that assurances about OpenAI’s nonprofit structure were “the hook for Altman’s long con.”

ppsreejith··on Batteries: How cheap can they get?
I think the unit is off. Starting from 2410 GWh & a compound increase of 59% per year gives us: 61,915 GWh (2410 * 1.59^7) which is about 61.915 TWh. So perhaps the author meant 61.915 TWh instead of GWh.

No way is this in anyway close to 8 doublings though. That would take 12 years or by 2035. (1.59^12 = 261x)

ppsreejith··on LLMs aren't "trained on the internet" anymore
Very interesting prompt regarding the boxer Eupolus of Thessaly. Gemini Advanced gets this wrong as well as Llama 3 70b (as run on groq.com).

However, if I start with: "What is the earliest record of cheating in the olympic games?" Then all models get the question right. It's surprising that GPT-4 gets it right on the first go.

ppsreejith··on Underwater volcano eruption 7,300 years ago is the largest in recorded history
From https://www.reuters.com/article/idUSL1N2XV1HA/

> According to the U.S. Geological Survey, published scientific estimates of the global CO2 emissions for all on land and submarine volcanos “lie in a range from 0.13 gigaton to 0.44 gigaton per year.”

> This is a fraction of the CO2 produced by human activity. In 2021, the global CO2 emissions from energy combustion and industrial processes alone reached a record high of 36.3 billion tonnes (or gigatons, GT), data from the International Energy Agency (IEA) showed

ppsreejith··on Show HN: Real-time image generation with SDXL Lightning
IIRC Dashtoon studio allows you to create comics with consistent characters using stable diffusion: https://dashtoon.com/create
ppsreejith··on Groq runs Mixtral 8x7B-32k with 500 T/s
Thanks again! Hope I'm not overwhelming but one more question: Are you decoding with batch size = 1 or is it more?
ppsreejith··on Groq runs Mixtral 8x7B-32k with 500 T/s
Thank you, that demo was insane!

Follow up (noob) question: Are you using a KV cache? That would significantly increase your memory requirements. Or are you forwarding the whole prompt for each auto-regressive pass?

ppsreejith··on Groq runs Mixtral 8x7B-32k with 500 T/s
Thank you for doing this AMA

1. How many GroqCards are you using to run the Demo?

2. Is there a newer version you're using which has more SRAM (since the one I see online only has 230MB)? Since this seems to be the number that will drive down your cost (to take advantage of batch processing, CMIIW!)

3. Can TTS pipelines be integrated with your stack? If so, we can truly have very low latency calls!

*Assuming you're using this: https://www.bittware.com/products/groq/

ppsreejith··on Groq runs Mixtral 8x7B-32k with 500 T/s
Relevant thread from 5 months ago: https://news.ycombinator.com/item?id=37469434

I'm achieving consistent 450+ tokens/sec for Mixtral 8x7b 32k and ~200 tps for Llama 2 70B-4k.

As an aside, seeing that this is built with flutter Web, perhaps a mobile app is coming soon?

ppsreejith··on Meta AI releases Code Llama 70B
If you go by what Zuck says, he calls this out in previous earnings reports and interviews[1]. It mainly boils down to 2 things:

1. Similar to other initiatives (mainly opencompute but also PyTorch, React etc), community improvements help them improve their own infra and helps attract talent.

2. Helping people create better content ultimately improves quality of content on their platforms (Both FoA & RL)

Sources:

[1]Interview with verge: https://www.theverge.com/23889057/mark-zuckerberg-meta-ai-el... . Search for "regulatory capture right now with AI"

> Zuck: ... And we believe that it’s generally positive to open-source a lot of our infrastructure for a few reasons. One is that we don’t have a cloud business, right? So it’s not like we’re selling access to the infrastructure, so giving it away is fine. And then, when we do give it away, we generally benefit from innovation from the ecosystem, and when other people adopt the stuff, it increases volume and drives down prices.

> Interviewer: Like PyTorch, for example?

> Zuck: When I was talking about driving down prices, I was thinking about stuff like Open Compute, where we open-sourced our server designs, and now the factories that are making those kinds of servers can generate way more of them because other companies like Amazon and others are ordering the same designs, that drives down the price for everyone, which is good.

ppsreejith··on Heat pumps, more than you wanted to know (2023)
> The energy stored in the oil is analogous to the "outside heat" for heat pumps.

This isn't correct. Here's a thought experiment to clarify:

If we burn X MJ of oil, we transfer X MJ of energy as heat to our house.

Otoh, if we set up a generator (say 40% efficient) which produces electricity from oil and dumps waste heat into our house, and use that to power a 300% efficient heat pump, we get 0.6X (generator waste heat) + 1.2X (300% * 0.4X) = 1.8X MJ of heat using the same X MJ of oil.

Thus,we get an additional 0.8X MJ of heat (80% extra heating) from the same X MJ of oil just by using a heat pump (& generator) compared to burning it. I.e we're using the energy stored inside the oil more efficiently.

ppsreejith··on Heat pumps, more than you wanted to know (2023)
This doesn't explain how they work though? I've found the following simplified model of how heat pumps work useful, and why they have > 100% efficiencies (Typically 300% - 400% or more compared to burning/ resistance heating which can only reach 100% efficiency).

Heat always flows from a high temperature to a low temperature. However, we want heat to go the opposite way (i.e. from the cold outside to our warm houses). There's a technique to do this. We use a gas (the refrigerant) to transfer heat.

We first expand the gas (which cools it) until it cools below the cold outside. We then bring it near the cold outside where it now starts absorbing heat until it matches the cold outside temperature. We then move the gas and compress it until it's temperature matches/exceeds our desired warm inside temperature. Then, we bring it to the warm indoors where heat will now flow out of it.

During the entire cycle, the gas is inside a closed loop. It exchanges heat through radiators. The compression & expansion cycles uses energy (technically they can be offset against each other a bit) which is added to the gas (conservation of energy) raising its temperature.

Thus by supplying a little electricity, we're able to move heat from the cold outdoors to the warm indoors. The high efficiencies are because for using X units of energy, we're able to heat the house by X+Y units where Y is the heat transferred from the outside. Typically Y >> X.

ppsreejith··on Heat pumps, more than you wanted to know (2023)
> It has efficiency over 100%, because you get all of the electricity's energy, plus some of the outside heat.

To add to this, the typical efficiency is 300% or more. Since typically most of the heat comes from being "transferred" from the outside.

ppsreejith··on Earliest Carpenters: 476k-year-old log structure discovered in Zambia
I like how this can be read both as a future where we regress technologically, or as a future where all work is in the domain of math.
ppsreejith··on Google's advanced music generation model and two new AI experiments
> You realise that this is a problematic situation created by technology

I'm saying this technology is a net positive for me. Sure, problems can arise from applying this net-positive technology but even considering them, they're a net positive. Especially in this case where we already have solutions for large amounts of crap content created by humans today. I think they'll scale well to crap content created by AI tomorrow (eg: Recommendation systems, p2p sharing etc).

Regarding your second point. I think we just have different perspectives on this. I see this technology as allowing me to do "strictly more" than I could before. You say that it's not "my expression" but I say it is because I'm able to bring forth what's in my mind (according to me!). You say it's not what I had in mind but compare against this scenario: I have an idea for a song and use a musical instrument. The output that I produce is quite far from my intent as I'm untrained. With this model, I can get a lot closer to my intent compared to using an instrument since I'm untrained. I still have the option of later learning an instrument if I'm unsatisfied.

Overall, I lose nothing and gain a nice power and still retain the option to learn an instrument if I wish. Thus it's a clear net positive for me that allows me to express more right now.

ppsreejith··on Google's advanced music generation model and two new AI experiments
Agree that training without permission and compensation is wrong. My argument is more that this tool is a net good. If suppose in the future, Someone releases a similar model trained on purely licensed data (Eg: Like Meta did with their Cm3leon model[0]), my arguments still stand that it gives people more choice and allows them to express more.

[0] https://ai.meta.com/blog/generative-ai-text-images-cm3leon/

> As such, and as described in our paper, we’ve trained CM3leon using a licensed dataset.

ppsreejith··on Google's advanced music generation model and two new AI experiments
Regarding the first point of being buried under an avalanche of crap, we can solve this problem through many ways. We can still filter by human music. Or we can use better recommendation systems (automated or p2p/word-of-mouth).

Regarding whether it is "my" expression or not, I consider expression as a way of making my feelings or thoughts known. This grants me an additional way of making my feelings or thoughts known. Perhaps of lower fidelity than if I learn to play an instrument but maybe that will improve with time. Either way, since I now have an additional choice of expressing my thoughts and _I_ decide if it's "good enough" to publish, I consider it as allowing me to express more.

ppsreejith··on Google's advanced music generation model and two new AI experiments
Not sure I agree with the negative opinions in this thread. Surely, having an additional option to listen to the song you want to is a good thing? It's not like the choice to listen to human artists is taken away. Furthermore, a lot more people can now express themselves through music that were previously unable to.
ppsreejith··on Why Children of Married Parents Do Better
Agree that divorce in extended families can be worse. I was making the (additional) claim that matrilocality & children being raised by woman's family can make families more stable (i.e divorce resistant).
ppsreejith··on Why Children of Married Parents Do Better
There's a lot of confounding variables here. Divorce could be a symptom of an underlying cause which could be the actual reason why kids do worse. Or it could be that when people pool their resources, it creates a more stable home (in resources and labor) and stable homes are a stronger cause.

Re:stable homes, it'd be interesting to envision whether non-nuclear family structures would be more stable. Maybe matrilineal/ matrilocal cultures (even if they're patriarchal) where children of a family are those born to it's women (more stable family since you don't divorce your siblings/mom, much freer relationships, paternity maybe unknown etc). The book Sex a At Dawn (controversial book!) explores these ideas.

ppsreejith··on The Smartest Person Who Ever Lived (2015)
For the given title, it's surprising that no non-Europeans are mentioned or non-STEM for that matter. Some other contenders could be:

1. Su Song

2. Plato

3. Gautama

4. Zhang Heng

5. Bhaskara(I/II)

6. Al Biruni

7. Da Vinci

8. Al Khwarizmi

ppsreejith··on Meta pitches EU to charge a $10 subscription for ad-free Facebook and Instagram
Meta doesn't sell your data to brokers. They sell access to your attention using ad slots. Advertisers upload their creatives to Meta and specify who they want to target.
ppsreejith··on Exllamav2: Inference library for running LLMs locally on consumer-class GPUs
IIUC this should work on the RTX 3090 as well (probably at less than 35tps)? Since the minimum requirement seems to be 24GB of VRAM
ppsreejith··on Google was founded 25 years ago today
Agreed on many counts but they did collude (with Apple, Intel, Intuit, Pixar, Lucasfilm allegedly due to bullying from Apple) to suppress poaching workers from each other. I'd say Facebook played a huge role in refusing to collude (as can be seen from this exchange between Sheryl Sandberg and Jonathan Rosenberg [1]) and thus driving up worker wages due to competition.

[1] https://techcrunch.com/2014/03/24/sheryl-sandberg-facebook-r...

[2] https://news.ycombinator.com/item?id=34227388 - Steve jobs email asking Google to stop poaching from apple and them agreeing

ppsreejith··on My Caste
Reading the attachments, the data seems to be recent (except for IIT bombay?) and includes assistant, associate, and tenured professors. Surprising that almost 175 of the 180 faculty members are from the "forward castes" (~30% of population) vs 5 (1 obc, 4 sc, scheduled caste) from the remaining 70%. Could also be due to academia skewing older? (i.e reflecting past biases)
ppsreejith··on The EU's war on behavioral advertising
Thanks, this does mirror my experience in many ways so I'll try to reply to each of your points. In my comment, I did talk about the value of advertising itself (not just contextual/behavioural).

Regarding your experience, a lot of advertising is retargeting or keeping the brand alive in the minds of buyers or remind them constantly (since humans are forgetful). This statistically increases the probability of purchases made amongst a cohort (very measurable). Personally, I think this is user-hostile and is a strong case for having more control over the kinds of ads we're exposed to or limit/penalise harmful ones. Taken to the extreme (going on a slight tangent here), if we really owned/controlled our devices and software, which means having control over consumption, we'd block all ads which would incentivise platforms to remove the distinction between content and ads and we'd do more content filtering on the client (This is however computationally inefficient, a sort of arms race).

Regarding influencing behaviour, it's a lot more efficient to match the right seller to the right buyer rather than influence buyer behaviour (which is hard). However, to achieve that extra marginal gains/returns on ad spend, advertisers often do attempt to change buyer behaviour (through building brand associations, retargeting etc). Again, better user control would help here.

Agree on street ads and other public real world ads as there is no consent here (Real world movement is a need, not a choice).

Regarding ads vs word-of-mouth, Ads really are very effective for many people, in a measurable way, both quantitative and qualitative. It of course will differ from person to person.

Regarding increased inefficiency and effect on the economy more generally, ads actually increase efficiency in the system as long as they are not rent-seeking (i.e. pay to jump the queue) or user-hostile. Eg: It helps many niche businesses reach customers which are otherwise very hard to reach (due to geographical or novelty constraints). Taken to the extreme, all value gained from advertising will ultimately flow to the ad platforms. Even in this extreme case, overall consumption increases (due to buyers buying more) as well as competition (due to many more competitive small businesses, innovators dilemma theory) which is good for a capitalist economy (i.e. makes it more dynamic in both choice and inequality) and increases GDP (a poor metric).

As an aside, it's also a progressive tax (I.e rich and poor both consume the same good but advertisers pay more to reach the rich, thus funding the good for the poor).

Some extreme opinions below on the fundamentals:

<Rant>

Even in localist/anarchist/decentralised utopias, ads won't die (especially personalised). It's a fundamental need of humans and the only other form of payment (apart from paying for compute itself) that doesn't depend on a system of violence to enforce it. As long as there exists economies of scale, there will be marketing (to increase consumption and hence cheapening it per capita i.e. use less labor), and hence advertising (to solve the unknown-unknown problem or presenting/impressing).

i.e. Even when we eliminate competitive enterprise, build localised production of needs, and have full control of our devices, if we want to build something cheaply or spread an idea, we'd have to advertise. It would look a lot more like "advertising on the merits" with full user control though considering human nature, there will also be a lot of "advertising focused on presentation".

</Rant>

← PreviousPage 2 of 4Next →