HNHacker News
TopNewBestAskShowJobs

ersatz_username

87 karma · joined January 27, 2021

submissionscomments
ersatz_username··on We're hosting Qwen3.8-27B for Free API usage
Hey HN,

The new Qwen3.8 release is really great and I know many people may not have GPUs that allow them to easily test it out so we are making it free to use through at least through the rest of the month.

ersatz_username··on Show HN: Playground to Test Skills and MCP
Hey Matt, would this allow me to centralize skills across different providers? Like rather than having a separate set of skills files for anthropic and openai they are just exposed via mcp?
ersatz_username··on Data shows public AI repos may be quietly becoming a supply chain risk
We analyzed over 1.8 million Hugging Face model repositories and found widespread licensing ambiguity, risky serialization formats, and subtle file-level inconsistencies—including drift between declared and actual artifact content. Even among the most-downloaded models, a surprising number are missing licenses or contain flagged files. Curious how others are thinking about model integrity and compliance in production.
ersatz_username··on We are not empty: The concept of the atomic void is a mistake
The word “empty” carries no semantic content under this construction of the term. Terms like “empty” are always used contextually, hence, why only a fool would correct their partner for mentioning the empty refrigerator by remarking on the presence of air.

Unless you think Sagan would have been surprised by the presence of an electric field (mediated by virtual particles…) between electrons and protons in an atom it’s quite likely you’re choosing an obtuse understanding of the term.

ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
Weird, sorry about this! We just removed the search and redeployed the docs. Hopefully that will fix it until we can sort the problem out. Would you mind giving it another shot?
ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
What does your data stack look like? I'll put something together specifically for you.
ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
A lot of these tools are very, very different from each other so it's hard to address each individually. Just by way of example, databend is a full on datawarehouse while greatexpectations is a testing framework evaluating data assertions (i.e. "I see there are nulls, but you wrote a test which says there shouldn't be").

Here are some things we think are really important though

1. Data quality testing ideally happens during CI not after merge.

2. Developers come first. Virtually every aspect of the tool can be customized, modified, and extended down to the basic data model without changing any upstream core code. Want to build your own custom application on top of your data lineage? Great! Have at it!

3. Users should be able to own not just their own data but their own metadata. We go to great lengths to maintain feature parity between the cloud and self-hosted application.

ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
I wouldn't personally draw such a bright line between monitoring and reliable CI/CD. That division definitely exists but partly as a product of the complexity introduced by fragmented data systems. In some ways an ideal world is one where the need for extraordinarily complex monitoring tools is actually pretty limited because we had tools to validate end to end data pipelines before making code changes if that makes sense.

We actually already do data monitoring as well although we haven't built the specific alerting features of Monte Carlo. There are quite a few tools that do that really well so it's not our focus at the moment.

ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
It depends a bit on your stack. Out of the box it does a lot with the metadata produced by the tools your using. With something like dbt we can do things like extract your test assertions while for postgres we might use database constraints.

More generally we can embed the transformation logic of each stage of your data pipelines into the edge between nodes (like two columns). Like you said, in the case of SQL there are lots of ways to statically analyze that pipeline but it becomes much more complicated with something like pure python.

As an intermediate solution you can manually curate data contracts or assertions about application behavior into Grai but these inevitably fall out of sync with the code.

Airflow has a really great API for exposing task level lineage but we've held off integrating it because we weren't sure how to convert that into robust column or field level lineage as well. How are y'all handling testing / observability at the moment?

ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
Really appreciate the kind words. If you don't mind my asking, what sort of issues were y'all experiencing that prompted you to start looking for solutions now?
ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
We aren't open source because we want to get anything out of it is the short answer. Of course to each their own but I've personally gotten a ton of value from open core tools in the past.
ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
Totally fair and appreciate the (well written) thoughts.
ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
Thanks so much! Really appreciate the kind words.

We haven't had anyone request Gitlab yet but would love to add support! Any chance you'd be willing to beta test for us? If so, shoot me an email at ian@grai.io :).

EDIT: It looks like the index issue is related to our search provider. Were you able to eventually load the page or is it fully blocking you?

ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
Just to be clear, the only limitation imposed by the license is preventing someone from reselling a cloud hosted copy of the tool. The code is otherwise totally free to use fork / modify / etc...
ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
We are pretty open to feedback on licensing and have gone back and forth internally because, frankly, we'd rather use a copy-left license.

We believe a project like this needs financial backing and a dedicated team driving development along but therein lies the tension. The common monetization paths either feature-lock critical self-hosted capabilities like SSO behind a paywall and/or monetize behind a cloud hosted option.

The Elastic license is an attempt to maintain feature parity between the cloud and self-hosted tool while still being protected from something like the big cloud providers ripping the code off altogether.

In all seriousness though, we would love to hear suggestions if you think there's a better path.

ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
Pre-built integrations are a big part of what makes onboarding easy but it sort of ends up in a catch-22 situation where whichever integrations gets highlighted is only directly applicable to the people using those tools.

If you have a different toolset onboarding will look exactly the same though, there's nothing truly DBT specific at work here. It's a good idea though! We really should put together a few other combinations so more people can see their own stack represented.

ersatz_username··on Launch HN: Grai (YC S22) – Open-Source Data Observability Platform
Shoot! I wasn't familiar with DarkReader but I just created a ticket to see if we can get it fixed. We recently redid the website and there's still plenty of room for improvement. Thanks for pointing that out :).
ersatz_username··on Launch HN: Windmill (YC S22) – Turn scripts into internal apps and workflows
Congrats on the launch Ruben! It's really exciting to see an open source tool like this out in the wild.
ersatz_username··on U.S. Debt ceiling can be sidestepped by minting a $1T platinum coin
How many times was the debt limit raised or suspended under the last republican administration (or the one before that, or the one before that…).

Or how about this one, who raised the debt limit more? Reagan or Obama? Fun fact, that would be Reagan.

Perhaps this isn’t a partisan issue.

ersatz_username··on Show HN: DataProfiler – What's in your data? Extract schema, stats and entities
This is really cool! I'm glad there's more work going into the area - visions[1] is a similar tool which embeds the type system (user defined schema) into a traversable graph rather than encoding all type information in the trained network. We originally wrote it as the backend for pandas-profiling[2].

I'm not sure if any of the authors are on here but y'all might look into Sherlock[3] as well. They've got pre-trained models for many other semantic types than you've currently implemented.

1. https://github.com/dylan-profiler/visions

2. https://github.com/pandas-profiling/pandas-profiling

3. https://sherlock.media.mit.edu/

ersatz_username··on Overnight Pizza and the Consistent Unreliability of Expert Guidelines
Bad analogy, seatbelts aren’t intended to operate in the course of normal driving rather their effect is observed during accidents. It’s a weird nuance but easier to see if you take a step back - it’s obvious why “I take a shower everyday and I’ve never needed a seatbelt” is flawed, namely, there’s no causal linkage between showering and seatbelts. The same is true for driving in this case, driving may be a necessary condition for an accident but it’s not sufficient.

To put it another way, your sample size for the test case of seatbelt value, i.e. an accident, appears to be zero rather than large as stated originally.

ersatz_username··on Show HN: Salary Standoff – Resolve the salary expectation/range stalemate
It’s really not intended for negotiations, it’s intended to determine if negotiation is possible. Candidate puts in their minimum, recruiter puts in their max, if max > min then negotiation is possible.

If the recruiter was able to make multiple submissions they could find the candidates minimum expectation exactly.

ersatz_username··on Euler's Fizzbuzz (2020)
You’ve got a factor of three sneaking into your 4^4.

In general (ab) mod n == (a mod n)(b mod n) mod n

In the case of (3*c)^4, 3^4 mod 15 -> 108 mod 15 -> 3.

ersatz_username··on You can't censor away extremism or any other problem
Science, as colloquially understood, seeks to empirically validate truth - true things can cause “harm.” So which is it? Are you looking for censorship of untrue ideas as discovered through open scientific inquiry, or, are you looking to censor ideas which cause harm. You can’t have both.

Put that aside, first principles, how can you hold a robust scientific debate on a topic that’s censored? The historical and common sense evidence strongly indicates it’s not possible.