HNHacker News
TopNewBestAskShowJobs

SnowflakeOnIce

511 karma · joined February 7, 2018

submissionscomments
SnowflakeOnIce··on Beyond A*: Better Planning with Transformers
They don't report execution time in the paper. It's likely that the A* implementation would run faster in terms of CPU or wall clock.
SnowflakeOnIce··on Beyond A*: Better Planning with Transformers
Note that the paper doesn't measure execution time of the A* or Transformer-based solutions; it compares length of A* algorithm traces with length of traces generated by a model trained on A* execution traces.

I suspect that the actual execution time would be far faster with the A* implementation than any of the transformers.

SnowflakeOnIce··on Magika: AI powered fast and efficient file type identification
I have seen 'file' misclassify many things when running it at large scale (millions of files) from a hodgepodge of sources. Unrelated types getting called 'GPG Private Keys', for example.

For textual data types, 'file' gets confused often, or doesn't give a precise type. GitHub's 'linguist' [1] tool does much better here, but is structured in such a way that it is difficult to call it on an arbitrary file or bytestring that doesn't reside in a git repo.

I'd love to have a classification tool that can more granularly classify textual files! It may not be Magika _today_ since it only supports 116-something types. For this use case, an ML-based approach will be more successful than an approach based solely on handwritten heuristic rules. I'm excited to see where this goes.

SnowflakeOnIce··on Magika: AI powered fast and efficient file type identification
Yes!

Sometimes a file has no extension. Other times the extension is a lie. Still other times, you may be dealing with an unnamed bytestring and wish to know what kind of content it is.

This last case happens quite a lot in Nosey Parker [1], a detector of secrets in textual data. There, it is possible to come across unnamed files in Git history, and it would be useful to the user to still indicate what type of file it seems to be.

I added file type detection based on libmagic to Nosey Parker a while back, but it's not compiled in by default because libmagic is slow and complicates the build process. Also, libmagic is implemented as a large C library whose primary job is parsing, which makes the security side of me jittery.

I will likely add enabled-by-default filetype detection to Nosey Parker using Magika's ONNX model.

[1] https://github.com/praetorian-inc/noseyparker

SnowflakeOnIce··on An Introduction to the WARC File (2021)
As to the your second question: WARC is based on an older format that was started in the 1990s, before SQLite existed.
SnowflakeOnIce··on Memory safety is necessary, not sufficient
I think it will stay that way for the foreseeable future (but who can say). Ways to fix the particular hole:

(1) disable creating new `code` objects directly from Python. This probably would break lots of things.

(2) Add a bytecode verification mechanism that would reject `code` objects whose bytecode would result in memory errors when executed. This could be a lot of implementation work; I'm not sure.

SnowflakeOnIce··on Memory safety is necessary, not sufficient
> Conventionally, in software security, Python is considered a memory-safe language. The piece makes the case that Python isn't memory safe when you FFI into a C library.

Interesting and largely unknown trivia: it's possible to invoke memory errors in the underlying C interpreter from pure Python code — no libraries and no imports needed!

One way of doing this is by creating new `code` objects with crafted bytecode. There is no bytecode verifier in Python to make sure, say, referenced stack variables in the VM are valid...

SnowflakeOnIce··on NY Governor vetoes ban on noncompete clauses, waters down LLC transparency bill
I believe MA (which overhauled its noncompete laws a few years back) requires "garden leave" in the case of a noncompete — the employer would have to pay the employee's salary for the period of enforcement.
SnowflakeOnIce··on How big is YouTube?
Surely someone at YouTube could answer this question definitively?
SnowflakeOnIce··on GitHub: Can no longer search code without being logged in
> I’d like to see these many occurrences.

Some high-profile cases where credentials were leaked on public GitHub: Uber in 2014 and 2021 [1, 2] and Twitch in 2021.

> Search is still possible using google and other methods

Yes, you can search with Google or other sources. But the thing is, pretty much only GitHub has easy access to all the code present there, readily available to search within minutes of pushing.

You could try mirroring GitHub yourself, but you'd need enormous disk space and bandwidth, and would quickly hit rate limits. You also wouldn't be able to do fast full regex search like you can with GitHub's search, as you don't have their search infrastructure.

Hackers are aware of this and do make use of GitHub's search to identify possible leaked credentials. I recently experimented with uploading a dummy AWS credential pair to a public Git repo, and saw that numerous IPs started trying those credentials less than 5 minutes after I pushed.

> any theoretical gains from forcing login are dangerous to count on as vulnerable projects are still vulnerable

Indeed, requiring login is not going to _solve_ the problem. But it could lessen the impact of leaked credentials at large scale, by making it more difficult for automated systems to harvest them.

Perhaps more significantly, requiring login could give better audit trails in incidence response situations, as the logs would indicate which accounts were searching for secrets.

> I think the solution to projects with hard coded credentials are to make the credentials easier to find and exploit so they fixed more quickly after being created

Yes, this can help! There are several other companies that specialize in secret detection. I've also written Nosey Parker, a fast regex-based detection tool that has higher-precision rules than similar tools [4]. GitHub also has its own offering in Advanced Security to address this problem.

[1] https://www.reuters.com/article/uk-uber-tech-lyft-hacking-ex... [2] https://www.securonix.com/blog/securonix-threat-research-ube... [3] https://news.ycombinator.com/item?id=28770590 [4] https://github.com/praetorian-inc/noseyparker

SnowflakeOnIce··on GitHub: Can no longer search code without being logged in
Yes, you are correct. But requiring that one be logged in to use the search functionality makes a stronger audit trail possible: looking at the search logs would indicate who has been hunting for secrets!
SnowflakeOnIce··on GitHub: Can no longer search code without being logged in
Another factor: anonymous faceted regex search across a huge volume of code allows bad actors to find hardcoded credentials and gain access to additional systems, without a good audit trail.

But yes, there are multiple good explanations for why they would lock down the API.

SnowflakeOnIce··on GitHub: Can no longer search code without being logged in
Aside from performance concerns, there have been many occurrences of bad actors using code search to find hardcoded credentials, and then using that to gain unintended access to additional systems.

I suspect the change has more to do with security concerns (and having an audit trail) than performance.

SnowflakeOnIce··on Star observatories you can visit in the United States
The article is wrong. It's an hour northwest of Boston.
SnowflakeOnIce··on Textual Adds a Command Palette
This is not quite a command palette, but in macOS apps, you can press command-shift-/ to search the names of menu commands. I use that frequently in apps that I don't use regularly (and hence don't remember where all the menu items are). Very helpful for discoverability.
SnowflakeOnIce··on Do you know how much your computer can do in a second? (2015)
Depending on many factors (like details of the patterns used and the input), some regex engines (like Hyperscan) can match tens of gigabytes per second per core. Shockingly fast!
SnowflakeOnIce··on How Python virtual environments work
'pip freeze' will generate the requirements.txt for you, including all those transitive dependencies.

It's still not great though, since that only pins version numbers, and not hashes.

You probably don't want to manually generate requirements.txt. Instead, list your project's immediate dependencies in the setup.cfg/setup.py file, install that in a venv, and then 'pip freeze' to get a requirements.txt file. To recreate this in a new system, create a venv there, and then 'pip install -c requirements.txt YOUR_PACKAGE'.

The whole thing is pretty finicky.

SnowflakeOnIce··on Running LLaMA 7B on a 64GB M2 MacBook Pro with Llama.cpp
The C API for Python changes between major Python versions. This means that native-code extensions (like torch) need to be build specifically for each version of Python they want to support.
SnowflakeOnIce··on Running LLaMA 7B on a 64GB M2 MacBook Pro with Llama.cpp
It's more an issue of torch not yet providing prebuilt binary wheels for Python 3.11. You could probably get it working, but it would involve building torch from source, which can be rather more involved than 'pip install'.

It's a packaging/release engineering limitation more than a language incompatibility thing.

SnowflakeOnIce··on Show HN: Nosey Parker, a fast and low-noise secrets detector for textual data
Yes and no.

On the one hand, Nosey Parker is effectively a special-purpose `grep` with a bunch of security-relevant patterns built-in, including one for PEM-encoded keys: <https://github.com/praetorian-inc/noseyparker/blob/main/data...>

On the other hand, to naively run the check you describe, you would need access to a copy of all of GitHub, which isn't feasible.

What you can do with Nosey Parker is use its GitHub enumeration features to specify your GitHub organization and a list of GitHub usernames you are interested in, and scan against just those. This will implicitly list all the relevant public repositories, clone them, and scan their entire history.

For your use case, another thing you could do is use the new GitHub code search (<https://cs.github.com>) to regex search for particular keys or tokens. That new search seems to cover lots of the public content available on GitHub.

Also, to put some color on this use case: in offensive security engagements (aka "red team" engagements) at Praetorian, we frequently find leaked credentials or tokens on GitHub or elsewhere, which allow us deeper access into the client's systems. It's a significant problem.

SnowflakeOnIce··on Show HN: Nosey Parker, a fast and low-noise secrets detector for textual data
That's an interesting use case!

The development so far has focused on security-related uses, especially finding hardcoded credentials in source code and log files. For the serial number and license spelunking you describe, Nosey Parker would need a couple more pieces to work well:

(1) Regex rules that would match serial numbers. I think it would be tough to write a regex for this that would match precisely and with good recall, i.e., producing few false positives and few false negatives. Not a showstopper, but not ideal.

(2) Built-in unarchiving support. Nosey Parker will handle textual files on disk just fine, but won't do useful things for many binary file types such as zip files. Your old external drives probably have tons of non-textual files like that. I would like to eventually add built-in unarchiving support to Nosey Parker; this would allow a way for textual content to be extracted from binary files and then scanned like the rest.

SnowflakeOnIce··on NP-complete isn't always hard
This reasoning works okay until you have to deal with machine-generated inputs, at which point your app freezes when it gets sent down a quadratic or worse complexity path.
SnowflakeOnIce··on Rust vs. Haskell
I understand your point and perhaps I'm in fierce agreement.

OTOH, things like parser combinators are much more ergonomic in Haskell than Rust.

SnowflakeOnIce··on DuckDB – An in-process SQL OLAP database management system
My experience with a modest machine and a ~100GB dataset was that DuckDB was significantly easier to use and much faster (20x) than Ray Data. Have not compared directly with Spark.

There was no cluster to set up or administer using DuckDB.

SnowflakeOnIce··on DuckDB – An in-process SQL OLAP database management system
DuckDB is terrific. I'm bullish on its potential for simplifying many big data pipelines. Particularly, it's plausible that DuckDB + Parquet could be used on a large SMP machine (32+ cores and 128GB+ memory) to deal with data munging for 100s of gigabytes to several terabytes, all from SQL, without dealing with Hadoop, Spark, Ray, etc.

I have successfully used DuckDB like above for preparing an ML dataset from about 100GB of input.

DuckDB is undergoing rapid development these days. There have been format-breaking changes and bugs that could lose data. I would not yet trust DuckDB for long-term storage or archival purposes. Parquet is a better choice for that.

SnowflakeOnIce··on Show HN: I trained an AI model on 120M+ songs from iTunes
What vector database do you use? Did you run into any scalability challenges there?
SnowflakeOnIce··on Google's OSS-Fuzz expands fuzz-reward program
Very interesting!

> Strangely, these bugs were found by the CI of ClickHouse, and not by any of the hundreds of other products using these libraries.

A lot of software projects that have "OSS Fuzz integration" have very naive or superficial fuzz harnesses that couldn't ever get thorough coverage.

If the code is never executed by the fuzzer, bugs won't be found by fuzzing.

In your case, fuzzing your application code, you likely end up executing code in those dependencies that their own fuzzers miss.

SnowflakeOnIce··on How do you know when macOS detects and remediates malware?
You get a warning popup and the application is blocked from running if it is not signed by an Apple Developer Account ($100/yr) and countersigned (i.e., notarized) by Apple.

This is separate from the 30% App Store commission.

SnowflakeOnIce··on Type-checked keypaths in Rust
I had not heard of keypaths before. They seem to provide functionality similar to lenses in Haskell.
SnowflakeOnIce··on Construction is life
zlib went more than 5 years without a release, and is heavily used.
← PreviousPage 2 of 4Next →