Searchcode.com’s SQLite database is probably 6 terabytes bigger than yours
boyter.org
boyter.org
It's an organization with an unpredictable return on investment, in practice, they don't really have any negative consequences if they waste public money, or if it was actually useless (unless too obvious to external people).
It's somewhat part of investing into experimental science.
In the eyes of the taxpayers, there's not that much of a difference between particle physicists who have been spending billions of dollars on bigger and bigger colliders, and psychologists who popularized non-replicating studies based on a sample of 35 university students but refused to touch IQ (one of the most, if not replicated psychology concepts ever), or biologists who invented new species out of thin air based on the most minuscule differences between population just to stop a development project they oppose for ideological reasons. All of this will cause severe backlash. Baby will be thrown out with the bathwater, trust and funding for science as a whole will suffer for decades.
> The current US Presidential administration is in the process of throwing the entire bathroom out, tub and all.
Are they, maybe, I have not seen any popularized document detailing their plans and they are taking large steps that will have consequences. Since there is not well popularized plan uncertainty and doubt is filling that void, a consequence which is easy to anticipate(as are some consequences of uncertainty and doubt).
> We're going to have substantially less reliable data over the next decade.
With no well popularized plan, hard to tell, and even harder because there are people hard at work on their own vision of the future that do not 100% line up with the current administration's plans and those people will influence the next decade as well.
National funding agencies are fundamentally in the business of choosing "this not that" because of the constraints of finite funding. Just like with VC math, it's difficult to imagine the benefits of the billion dollar collider outweigh the opportunity cost of investing in a large nation's worth of researchers to explore smaller, cheaper ideas.
I've seen her writings here and there over the years, and I remember she was generally respected as an opinionated expert. What happened?
--
[0] - I've seen my share of dramas which started with a message like this video, and where I thought I had a good picture, until some time later some critical details came to light and I ended up flipping my stance on it 180°.
Respected and opinionated experts gotta eat
(I'm not just reacting to the existence of a sponsorship section alone - more to the choice of the sponsor, the presentation, and the very juxtaposition of a "trust me while I quote from private conversation" and "this video was sponsored by a data broker company".)
You can find ways to dismiss opinions you don't like, but targeting individual contributors that use sponsorships to maintain or alleviate financial stability is an interesting attack. I guess if that's your only critique, you have no claims to any other negative critique about her points? Just a literal ad hominem? "What happened" to actual discourse?
It would be merely off-putting if it was tacked onto a video about hard, verifiable facts - but to make a video that explicitly asks the audience to trust her on her world, and then end it by saying the video itself was sponsored by a company? Surely she's aware of the optics? Even the regular YouTube influencers usually know better and skip the sponsorship section when making a complaint video.
I'm also not committing an ad hominem here (if anything, maybe some adjacent fallacy). This is the lens through which I view all YouTube channels, and it applies here, and doubly so given that this is not a video about independently verifiable facts - it's all based about an e-mail she claims she received, and she acknowledges that directly. We have to trust her at her word.
Now, I don't know about you, but if someone in one moment tells me some information, and then in the next moment starts giving me "capitalist brand advertisement speak" about some dubious product, I'm going to take it as a sign that someone doesn't actually care about my well-being, as manipulating someone into a bad deal for profit is plain malice[0]. Additionally, I might question that person's integrity - depending on obviously they're knowingly pushing a bad product, or how indifferent they are to what they're promoting. Which, in turn, will make me question everything they just told me before - after all, if they just demonstrated they're fine with lies or bullshit now, why should I assume they held themselves to the highest standards of scientific and personal truth moments before?
I'm honestly done complaining over YouTube creators; I just accepted that many of the well-known pop-sci channels turned into content marketing schemes. At least I stopped being surprised by exaggerations and inaccuracies in the main parts. I commented this time only because I totally didn't expect to see a reputable scientist doing this kind of stuff.
--
[0] - Yes, I stand firm on this, and yes, if you scale this view up, you end up considering the entire field of "marketing communication" (covering a subset of intersection of sales, marketing and advertising) as a cancer on modern society. I wrote an article about this some time ago, which I've been told has apparently popped up on HN last week.
> Just a literal ad hominem?
I do not read what in the comment as an attack, but an outsider to the field trying to judge the opinion they are absorbing.
Some people are going to expect an experts to have an outside money stream that comes from their expertise so in video adds are not necessary. That is not a heuristic that is going to be right 100% of the time, but is is not the worst heuristic or an attack.
I'd say, my own emotions aside, she made a pragmatic mistake here - first delivering a quite powerful bomb that she explicitly acknowledges we need to trust her about (as it's an excerpt of a private conversation, not possible to independently verify), only to tell us - also explicitly - that the video was sponsored by a company. This is the kind of stuff you find in late-stage capitalism jokes, or movies about corporate utopia where this juxtaposition is exaggerated for effect. I did not see it coming, and I struggle to understand why it did.
But more generally, while I wouldn't describe it as "targeting individual contributors that use sponsorships to maintain or alleviate financial stability", I do believe that what kinds of sponsorships people choose and how they fulfill the sponsor's conditions does directly reflect on trustworthiness of the entire message - to think otherwise would be to believe that humans are capable of compartmentalizing their activities into high-integrity and low-integrity parts, which is something I don't believe humans are capable of over long-term. Maybe I'm wrong about that - if psychology says it's normal, then perhaps I need to reconsider my heuristic. But if I'm right, then this is directly on-topic and I believe it's right to bring it up, to the extent the creator/speaker is asking the audience to take them on trust. Aka. on authority. Which means it applies doubly so to the experts - their choice of what and how they advertise matters more, because they're lending their credibility to both the message and the ad that pays for it.
I agree here. I would say historically this was a more common view. What I observe is that more people are disregarding this and it is in part due to feedback loops between the very large audiences that web platforms provide today and the creator. If you message connects with a large enough group of people with enough loyalty/trust(that say weird sponsor messages do not effect there viewership) you can safely disregard the rest of audience to a large degree. Delivering with emotions, "delivering a quite powerful bomb"s, etc help build that loyal group of followers but also lead to a feedback cycle that can make things more one side/hyperbolic/etc.
This has a knock on effect in people, at least those like me, who now devalue many similar emotional pleas without evidence for both good and ill. After all if people are incentivized to be delivering impassioned speeches and "bomb"s, statistically there are going to be more people who do so inappropriately, sometimes it seems somewhat normalized just due to how much I see it. To me the message she delivered had little to no impact on me because it was not delivered with either reasonable evidence or at least a start of a plan or solution to the issue she is describing. This knock on effect just makes things worst to some extent though since creators already in that feedback loop have even less incentive to reach out to someone like me because they have to overcome the additional barrier, and would not when the same level of loyalty even if they did.
> But if I'm right, then this is directly on-topic and I believe it's right to bring it up, to the extent the creator/speaker is asking the audience to take them on trust. Aka. on authority. Which means it applies doubly so to the experts - their choice of what and how they advertise matters more, because they're lending their credibility to both the message and the ad that pays for it.
I agree it is on topic and relevant, but it did not register to me at all in this video because the trust was lost when there was only the impassioned deliver without hard evidence or an action plan to help address the issue. It likely also does not register to those in the loyal impassioned group of followers either.
Independent of potential regulation, creating more high trust, potentially collaborative, sources of information seems like like a partial way to counter balance some of this.
Yes, science has extremely unpredictable return on investment.
What's your suggestion? Don't try?
Since I do a lot of cross-compiling I have been using https://modernc.org/sqlite for SQLite. Does anyone have some knowledge/experience/observations to share on concurrency and SQLite in Go?
This guy is great: https://fractaledmind.github.io/2024/04/15/sqlite-on-rails-t...
https://colin-scott.github.io/personal_website/research/inte...
(This version is more useful since it shows change over time)
Something something, Cunningham's Law.
This has been a big issue I have found when researching SQLite and trying to determine if it's suitable. I can't figure out the right way to do certain stuff or docs are difficult/outdated etc in end I always end up defaulting back to Postgres
Three modes are supported: one where the library does no locking and you can only use the library from one thread at a time; one where global structures are locked but not connections so you can use each connection from one thread at a time; and one where everything is locked.
For actual concurrency you probably want one database connection per thread.
My sole reference: https://www.sqlite.org/lockingv3.html
You should not share sqlite connections between threads (or anything even remotely resembling threads): while the serialized mode ensures you won't corrupt the database, there is still per-connection state which may / will cause issues eventually e.g. [1][2][3]. Note that [1] does apply to read-only connections.
You can use separate connections concurrently. If you're using transactions and WAL mode you should have a pool of read-only connections (growable) and a single read/write connection. If you're not using multi-statement transactions (autocommit mode) then you can just have a pool.
[1] https://sqlite.org/c3ref/errcode.html
[2] https://sqlite.org/c3ref/last_insert_rowid.html (less of an issue now that sqlite has RETURNING)
[3] https://sqlite.org/c3ref/changes.html (possibly same)
Go database/sql handles ensuring each actual database connection is only by a single goroutine at a time (before being put back into the pool for reuse), and no sane SQLite driver should have issues with this.
If you're working in Go with database/sql you're meant to not have to worry about goroutines or mutexes. The API you're using (database/sql) is goroutine safe (unless the driver author really messed things up).
To be clear: each database/sql "connection" is actually a "pool of connections", that are already handed out (and returned) in goroutine safe way. The advise is to use two connection pools: a read-only one of size N, and a read-write one of size 1.
Regardless, mutexes are not needed.
All sane Go SQLite drivers should be concurrency safe.
That said, (and I'm biased but) there's something fishy about modernc concurrency: https://github.com/cvilsmeier/go-sqlite-bench#concurrent
Querying data with 2 threads shouldn't be 2x slower than using 1 thread; 4 threads shouldn't be 6x slower; or 8 threads 15x.
The folks at GoToSocial have long maitained a conccurency fix; I don't know if it's related to this performance issue.
Still the advice about using read-only and read-write connections still holds, for all drivers.
Why? Because in SQLite reads can be concurrent, but writes are exclusive. If you're using WAL, a single write can be concurrent with many reads, otherwise a writer also blocks readers. And SQLite handles all of this.
But SQLite locks are polling locks: a connection tries to acquire a lock; if it can't, it sleeps for a bit and tries again; rinse, repeat. When the other connection finally releases its lock, everyone else is sleeping. There's a lot of sleeping involved; also with exponential backoff.
Using read-only and read-write connections "fixes" this. Make the read-only a pool of N connections, and the read-write a single connection. Then, all waiting will be in Go channels, which is fast at waking waiters, and doesn't involve needless sleeping/pooling. As a bonus, context cancellation will be respected for this kind of blocking.
On this final point, my driver goes to great lengths to respect context cancellation, both of CPU queries (which you need to forcefully interrupt), and busy waiting on locks (which also doesn't work for most other drivers). So you can set a large busy timeout (one minute is the default) with the confidence that cancellation will work regardless. The dual connection strategy still offers performance benefits, though.
So... you can get your ducks in a row in terms of checking sqlite docs, your sqlite config and compile-time options, and your runtime environment, and then remove the mutexes (modernc is another variable, but I don't know about that). Or, if it's working for you, you can leave them in.
Reasons to remove them might be... it's a pain to maintain them; it might help performance (sqlite is inherently one writer multiple readers, and since you're using RW locks, your locks may already align with sqlite's ability to use concurrency). If those aren't issues for you then you can leave them in.
> Making use of sql.DB from multiple goroutines is safe, since Go manages its own internal mutex.
Then I quote the database/sql documentation [2]:
> DB is a database handle representing a pool of zero or more underlying connections. It's safe for concurrent use by multiple goroutines.
[1]: https://stigsen.dev/notes/go-database-thread-safety/ [2]: https://pkg.go.dev/database/sql@go1.23.3#DB
Works like a charm, as in: the web app consuming the API linked to it returns paginated results for any relevant search term within a second or so, for a handful of concurrent users.
I've had decent success with `sqlite-zstd`[0] which is row-level compression but only on small (~10GB) databases. No reason why it couldn't work for bigger DBs though.
Additionally Xtranormal missed out on the generative video curve
A whole lot of people have already taken that advice.
What do you run this on? Just some aws vpc with a huge disk attached?
They do have a higher price floor, though. There are no $5/month dedicated servers anywhere - the cheapest is more like $40. There are $5/month virtual servers outside of AWS which are cheaper and more powerful than $5/month AWS instances.
That's neat! I bet it keeps growing a WAL file while the backup is ongoing right?
You may find sqlite_rsync better.
A reliable backup will need to use the .backup command, the .dump command, the backup API[0], or filesystem or volume snapshots to get an atomic snapshot.
``` dbRead, _ := connectSqliteDb("dbname.db") defer dbRead.Close() dbRead.SetMaxOpenConns(runtime.NumCPU())
dbRead, _ := connectSqliteDb("dbname.db") defer dbWrite.Close() dbWrite.SetMaxOpenConns(1)
```
dbWrite, _ := connectSqliteDb("dbname.db")
I was a bit confused by the code, then assumed it was a typo
But I liked his tip about SQLite driver scalability to avoid that stupid locked error that I too have faced regularly. numCPUS for readers and single writer - will try that out.
As a sanity check, "fwrite" only has 8 references in the entire database.
Yeah, agreed, I think the migration didn't actually work.
> By default searchcode prioritizes system survival (hey its a free service!), and as such might do some load shedding, which can mean you don't see results you expect. You can do some things to help with this.
I don’t know whether that explains the results. Some similar queries, e.g., searching for `math.ceiling`, do return many results.
It uses a simple crdt table structure, and allows me to have many live SQLite instances in all data enters.
A Nat Jetstream server is used as a core with all SQLite DBS connected to it.
It operates off the WAL and is simple to install.
It also replicates all blobs to S3 , with the directory structure in the SQLite db.
With a Cloudflare domain , the users request is automatically sent to the nearest Db.
So it replaces cloudscapes D1 system for free . Just a hetzner 4 euro cos is enough.
Mine is holding about 200 gb of data.
dbRead, _ := connectSqliteDb("dbname.db")
defer dbRead.Close()
dbRead.SetMaxOpenConns(runtime.NumCPU())
dbRead, _ := connectSqliteDb("dbname.db")
defer dbWrite.Close()
dbWrite.SetMaxOpenConns(1)
Is dbWrite ever declared? I know it's just an example, but still ...Now I have to ask myself whether this error resulted from relying on a human, or not relying on one.
I still chose postgres managed by AWS mainly to reduce the operational overhead, I keep thinking if I should have just gone with sqlite3 though
The author clearly mentioned that they want to work in the relational space. Choosing Mongo would require refactoring a significant part of the codebase.
Also, databases like MySQL, Postgres, and SQLite have neat indexing features that Mongo still lacks.
The author wanted to move away from a client-server database.
And finally, while wanting a 6.4TB single binary is wild, I presume that’s exactly what’s happening here. You couldn’t do that with Mongo.
Edit: ok the other guy mentioning mongo is clearly being sarcastic
6.4x($100+256x4x$1)=$7,193.6 worst case pricing
So this database would cost less than $8000 in DRAM chips hardware and can be searched/indexed in parallel in much less than a second.
Now you can buy old DDR2 chips harvested from ewaste at less than $1 per GB, so this database could cost as low as $1000 including the labour for harvesting.
You can make a wafer scale integration with SRAM that holds around 1 terabyte SRAM for $30K including mask set costs. This database would cost you $210.000. Note that SRAM is much faster than DDR DRAM.
You would need 15000/7 customers of this size database to cover the manufacturing of the wafers. I'm sure there are more than 2143 customers that would buy such a database.
Please invest in my company, I need 450,000,000 up front to manufacture these 1 terabyte wafers in a single batch. We will sprinkle in the million core small reconfigurable processors[1] in between the terabyte SRAM for free.
For the observant: we have around 50 trillion transistors[2] to play with on a 300mm wafer. Our secret sauce is a 2 transistor design making SRAM 5 times denser that Apple Silicon SRAM densities at 3nm on their M4.
Actually we would not need reconfigurable processors, we would intersperse Morphle Logic in between the SRAM. Now you could reprogram the search logic to do video compression/decompression or encryption by reconfiguring the logic pattern[3]. Total search of each byte would be a few hundred nanoseconds.
The IRAM paper from 1997 mentions 180nm. These wafers cost $750 today (including mask set and profit margin) and now you would need less than a million investment up front to start mass manufactoring. You just would need 3600 times the wafer amount compared to the latest 3nm transistor density on a wafer.
[1] Intelligent RAM (IRAM): chips that remember and compute (1997) https://sci-hub.ru/10.1109/ISSCC.1997.585348
[2] Merik Voswinkel - Smalltalk and Self Hardware https://www.youtube.com/watch?v=vbqKClBwFwI&t=262s
[3] Morphle Logic https://github.com/fiberhood/MorphleLogic/blob/main/README_M...
Not sure how you arrived at this calculation. 256x$4 already accounts for 256GB. The database in the OP is 6400GB large. So shouldn't it be 25x($100+256x$4) = $28100?
FWIW in practice this number should be much, MUCH lower. You can get 1TB of ECC DDR4 for ~$2k, probably lower if you buy wholesale.
The $450 million is also an estimate that could be off by an order of a magnitude.
My argument still stands: in-memory databases with SRAM wafers are much faster and cheaper than customers (managers) realize, no need to have database software.
In 1997 the IRAM paper was still thinking about 180nm DRAM chips, in 2025 we can do 3nm wafer scale integrations so Moores law gives us affordable Terabyte in-memory databases that cost less than a programmer to write and manage databases.
In 1981 Dan Ingalls wrote[1] in Design Principles Behind Smalltalk: "An operating system is a collection of things that don't fit into a language. There shouldn't be one."
IN 2025 I argue that "a database is a collection of tricks to shuffle data between compute and storage memory. There shouldn't be one.".
I would go one step further and say a von Neumann architecture or Harvard architecture microprocessor is a collection of gates to compute in separated memory storage. There shouldn't be one.
[1] https://www.cs.virginia.edu/~evans/cs655/readings/smalltalk....