HNHacker News
TopNewBestAskShowJobs

rusanu

996 karma · joined April 4, 2014

Robot at UiPath Former SQL Server engine dev at MS.
submissionscomments
rusanu··on Baffling ABC maths proof now has impenetrable 300-page ‘summary’
> “The language strikes me as substantially more accessible than that of the original papers.”

So is not really a summary, is more of a reformulation, an easier to read rewrite.

rusanu··on 'Frankenstein dinosaur' mystery solved
A new hypothesis of dinosaur relationships and early dinosaur evolution[0]

[0] http://www.nature.com/nature/journal/v543/n7646/full/nature2...

rusanu··on 'Frankenstein dinosaur' mystery solved
tl/dr: Theropods more closely related to ornithischians than to sauropods
rusanu··on Microsoft is Hiring Go engineers to work on Kubernetes
I also saw today this article Microsoft Prepares SQL Server 2017 for Linux and Containers [0], from this tweet[1] "Microsoft Prepares SQL Server 2017 for Linux and Containers, SQL Server lab runs on kubernetes and docker". If they moved the entire SQL lab (many thousands of machines) to Kubernetes is quite a big deal. They're probably warming up on how to manage it at scale, and I'm pretty sure that moving Azure SQL DB infrastructure to it is the goal. If they can get better isolation and higher density that today (some Azure Fabric derivative, afaik), it would save tonnes of moneys.

[0] https://thenewstack.io/sql-server-2017-brings-microsofts-dat...

[1] https://twitter.com/slava_oks/status/887359748047216640

rusanu··on StackOverflow for private teams in beta
Loose lips sink ... Nobel prizes? https://en.wikipedia.org/wiki/Photo_51
rusanu··on United Airlines piloting technology to manage the problem of oversold flights
Sometimes people miss flights, or last minute cancel, for various reasons. If past statistics show the every 200 seats plane has 10 no-show passengers at departure on average, you can sell 210 tickets in average and depart with every seat occupied. Hotels do it too, and many more industries.

https://en.wikipedia.org/wiki/Overselling

rusanu··on Microsoft is laying off thousands in a major global sales reorganization
probably Enterprise Agreements

https://en.wikipedia.org/wiki/Microsoft_Enterprise_Agreement

rusanu··on Delivering Billions of Messages Exactly Once
> it's still something that has to be dealt with somewhere

Database programmers have the means to deal with it off-the-shelf: BEGIN TRANSACTION ... COMMIT. When your queues are in the database, this becomes trivial. Even without the system I'm talking about (Service Broker) that has the queues stored in the database, most regular messaging systems do support enrolling into a distributed transaction and achieve an atomic dequeue/process sequence, is just that many apps/deployments don't bother to do it because the ops overhead (XA coordinator), reduced throughput and/or simply not understanding the consequences.

Point is that durable, persisted, transacted 'sockets' are behaving very differently from a TCP socket. Is a whole lot harder to simply lose a message in the app layer when interacting with a database.

rusanu··on Delivering Billions of Messages Exactly Once
My point is that I've seen people making decisions to go with 'best effort delivery' and live with the (costly) consequences because they read here and there that EOIO is impossible, so why bother trying.
rusanu··on Delivering Billions of Messages Exactly Once
Not when your 'socket' is a persisted, durable, transacted medium (ie. a database). Sure, applications can 'pull data out of the socket and then die', but this is a common scenarios on databases which is handled with transaction and post-crash recovery. The application comes back after the crash and find the same state as before the crash (the 'socket' still has the data ready to pull off), it pull again, process, and then commit. This is not duplicate delivery, since we're talking about an aborted and rolled back attempt, followed later by a successful processing. Again, databases and database apps have been dealing with this kind of problems for decades and know how handle them.

I've been living in this problem space for many years now and seen the wheel reinvented many times. Whenever the plumbing does not guarantees EOIO but the business demands it, it gets pushed into the app layer where TCP (retries and acks) is reimplemented, to various success levels.

rusanu··on Delivering Billions of Messages Exactly Once
Pat had some opinions about CAP and SOA and distributed systems, see [0]. I also remember a talk given by Pat and Eric Brewer together, that went deeper into the whole CAP ideas vis-a-vis the model Pat had been advocating (see Fiefdoms and Emissaries [1]), but I can't remember when it was or find a link for it.

[0] https://blogs.msdn.microsoft.com/pathelland/2007/05/20/soa-a...

[1] http://download.microsoft.com/documents/uk/msdn/architecture...

rusanu··on Delivering Billions of Messages Exactly Once
Yes, I'm talking about distributed systems and I am aware of the CAP theorem. Hence my choice of the word 'practical'.

As I said, users had cases when the plumbing (messaging system) recovered and delivered messages after +40 days of network partitioning. Correctly written apps completed the business process associated with those messages as normal, no special case. Humans can identify and fix outages and databases can easily outlast network outages (everything is durable, transacted, with HA/DR). And many business processes make perfect sense to resume/continue after the outage, even if it lasted for days.

rusanu··on Delivering Billions of Messages Exactly Once
Having spent 7 years of my life working with Pat Helland in implementing Exactly Once In Order messaging with SQL Server Service Broker[0] I can assure you that practical EOIO messaging is possible, exists, and works as advertised. Delivering data EOIO is not rocket science, TCP has been doing it for decades. Extending the TCP paradigms (basically retries and acks) to messaging is not hard if you buy into transacted persisted storage (= a database) for keeping undelivered messages (transmission queue) and storing received messages before application consumption (destination queue). Just ack after you commit locally.

We've been doing this in 2005 at +10k msgs/sec (1k payload), durable, transacted, fully encrypted, with no two phase commit, supporting long disconnects (I know for documented cases conversations that resumed and continued after +40 days of partner network disconnect).

Running into resource limits (basically out of disk space) is something the database community knows how to monitor, detect and prevent for decades now.

I really don't get why so many articles, blogs and comments claim this is not working or impossible or even hard. My team shipped this +12 years ago, is used by major deployments, technology is proven and little changed in the original protocol.

[0] https://docs.microsoft.com/en-us/sql/database-engine/configu...

rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
Can you clarify then wether the AV was simply not scanning the SQLite files (ie. the file extension or the file location was exempt in AV config) ?
rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
Since you mention the OR impedance mismatch problem, I have to link to The Vietnam of CS article: http://blogs.tedneward.com/post/the-vietnam-of-computer-scie...
rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
> Do a query, get back a handle to a file that you can treat just like you opened it yourself

Things are a bit more complex. For one, the trivial problem of client vs. server host. The DB cannot return a handle (a FD) from the server, because it has no meaning on the host running the app. The second problem is that any file manipulation must conform to the DB semantics for transactions, locking, rollback and recovery.

What you describe does exists, is the FileStream feature that dates back to 2007 if I remember correctly. I'm describing the SQL Server feature since this is what I'm familiar with. The app queries the DB for a token, using GET_FILESTREAM_TRANSACTION_CONTEXT[0] and then uses this token to get a Win32 handle for the 'file' using OpenSqlFilestream[1]. The result handle is valid for usual file handle operations (read, write, seek etc). There were great expectations on this feature, but in real life it flopped. For one it caused all sort of operational headache from the increased DB files size (increased backups size etc) or from problems like having to investigate 'filestream thumbstone status'[2]. But more importantly, adoption required application rewrite (to use the OpenSqlFilestream), which of course never materialized.

File Tables is a newer stab at this problem and this one does allow to expose the DB files as a network share and apps can create and manipulate files on this share and everything is backed by the DB behind the scenes. But turns out a lot of apps do all sort of crazy things with the files, like copy-rename and swap as means to do failure safe saves, but many such operations are significantly more expensive in DB context. And when the DB content is manipulated directly by the apps that 'think' they interact with the filesystem, a lot of useful metadata is never collected in the DB, since the file API used never requires it (think file author, subject etc).

[0] https://docs.microsoft.com/en-us/sql/t-sql/functions/get-fil... [1] https://docs.microsoft.com/en-us/sql/relational-databases/bl... [2] https://www.sqlskills.com/blogs/paul/filestream-garbage-coll...

rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
To name just a few:

Filestream https://docs.microsoft.com/en-us/sql/relational-databases/bl...

File Tables https://docs.microsoft.com/en-us/sql/relational-databases/bl...

Remote Blob Storage https://docs.microsoft.com/en-us/sql/relational-databases/bl...

BFILE http://docs.oracle.com/cd/E11882_01/appdev.112/e18294/adlob_...

I'm sure there are more. But rest assured, they do cost, and usually a lot.

rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
In this case nginx is the 'application'. the requests is still going to be expressed as a SQL query, sent to the PG, parsed, compiled, optimized, executed, then the tabular response formatted as the HTTP response. Many more steps compared to a file-on-disk response.

But I second that is an interesting nginx module

rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
Amen to that. Fastest database query is the one you never run.
rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
You have two consystency issues with storing the files in the filesystem:

- rollbacks in the DB can lead to orphaned files on disk. One can try to add logic in the app (eg. a catch block that removes the file if the DB rolled back) but that is not gonna help on a crash

- it is impossible to obtain a consistent backup of both the DB and the filesystem. You can backup the filesystem and the DB, but the two will not be consistent between them unless you froze the app during the backup. When you restore the two backups (filesystem, DB) you may encounter any anomaly: orphaned files (exists on filesystem but no entry in DB), broken links (entry in DB referencing a non-existent file) etc. This is because the moment at which the backup 'views' the file and the DB record referencing it are distinct in time.

As for write reordering: write-ahead log systems relies on correct write order. All DBs worth their name enforce this one way or another (via special API, via config requirements etc etc)

rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
But that would work only for static content, like the game assets. The discussion of BLOBs vs. filesystem comes up mostly in the context of content management and user/app uploaded content.
rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
From OP: "SQLite is much faster than direct writes to disk on Windows when anti-virus protection is turned on. Since anti-virus software is and should be on by default in Windows, that means that SQLite is generally much faster than direct disk writes on Windows."

I don't get this. If scanning the content is important (as acknowledged by the author), then bypassing the scan via blob storage is a security issue and the application should go through some extra hoops to scan the content before saving it to blob, and this should be measured and part of the comparison.

Also, if the SQLite files are exempt from AV scan, then the level field should also exempt the uploaded files folder in test. I mean, knowing the dice are loaded and then claiming it as an advantage does not seem professional.

rusanu··on SQLite small blob storage: 35% Faster Than the Filesystem
I must point out Jim Gray's paper To Blob or Not To Blob[0]. His team considered NTFS vs. SQL Server, but most rationale applies to any filesystem vs. database decision.

The summary was "The study indicates that if objects are larger than one megabyte on average, NTFS has a clear advantage over SQL Server. If the objects are under 256 kilobytes, the database has a clear advantage. Inside this range, it depends on how write intensive the workload is," but keep in mind this is spinning media from 2006. Modern SSDs change the equation quite a bit, as they are much more friendly to random IO and benefit less from database write-ahead log and buffer pool behavior.

Also, when deciding between blob vs. filesystem, blobs bring transactional and recovery consistency. The DB is self contained, and all blobs are contained in it. A restore of the DB on a different system yields a consistent system, it won't have links to missing files, and there won't be orphaned files left over (files not referenced by records in DB).

Despite all this, my practical experience is that filesystem is better than blobs for things like uploaded content, images, pngs and jps etc. Blobs bring additional overhead, require bigger DB storage (more expensive usually, think AWS RDS) and the increased size cascades in operational overhead (bigger backups, slower restore etc).

[0] https://www.microsoft.com/en-us/research/publication/to-blob...

rusanu··on Comdb2 – Bloomberg's distributed RDBMS under Apache 2
Probably the same motivation Yahoo had to release Hadoop, FB to release Hive, Netflix to release so much of their libs and so on and so forth:

- if nothing else, it does no harm (no 'secret sauce' competitors could benefit from)

- it buys karma (think recruiting goodwill)

If the project catches on though then there are many advantages:

- it can spark a self-sustained ecosystem that can further drive the product, at much lower cost for original creator (think Hadoop leading to Cloudera, Hortonworks etc). Product improves, bugs are fixed, toolset matures

- newhires come with know-how to use your internal tools, lower ramp up, better productivity. Anecdotal, but when I was at Microsoft no newhire knew how to use the internal Cosmos stuff, and even among old timers more folk were familiar with Hadoop...

rusanu··on A subway-style diagram of the major Roman roads, based on the Empire ca. 125 AD
Tangential: George Dow[0] and Harry Beck[1] created the 'Tube map'.

[0] https://en.wikipedia.org/wiki/George_Dow [1] https://en.wikipedia.org/wiki/Harry_Beck

rusanu··on British Airways IT chaos was caused by human error
> a contractor doing maintenance work inadvertently switched off the power supply [...] This resulted in the total immediate loss of power to the facility, bypassing the backup generators and batteries... After a few minutes of this shutdown, it was turned back on in an unplanned and uncontrolled fashion, which created physical damage to the systems and significantly exacerbated the problem.
rusanu··on Ask HN: What language-agnostic programming books should I read?
I second Code Complete
rusanu··on Ford to cut North America, Asia salaried workers by 10 percent
Is Oregon. I still remember my shock at the pump station after being told that I'm not qualified to handle explosive substances (gas)...
rusanu··on Spanner: Becoming a SQL System [pdf]
You can point back to the 2004 benefits overhaul [0], the famous Towels story [1], low compensation rates compared to Google, Amazon and Facebook [2]. Just go over minimsft.blogspot.com posts at the time.

Add to this the lack of vision and direction, catastrophic acquisitions, dismal flagship product releases. At the time there were running jokes about the Inbox filling up with "After 15 years, is time to send that email" subject lines...

  [0] http://old.seattletimes.com/html/businesstechnology/2001938654_microsoft26.html  
  [1] http://www.zdnet.com/article/microsoft-brings-back-the-towels-5000148135/  
  [2] http://minimsft.blogspot.com/2006/03/internal-microsoft-compensation.html
rusanu··on Spanner: Becoming a SQL System [pdf]
Tangential.

For me the fascinating thing is looking at the list of authors to recognize so many from the 2005-2012 Microsoft SQL Server team. Folk I know personally as exceptional performers. Same when I look at Aurora papers. I see this as the result of Ballmer's famous HR initiatives and the massive brain drain that occurred at Microsoft around 2010-ish.

← PreviousPage 2 of 6Next →