Source: I have lived in Saigon for a decade, which beats hands down the many books I had read about the subject before that
4,087 karma · joined August 3, 2016
Source: I have lived in Saigon for a decade, which beats hands down the many books I had read about the subject before that
https://www.ibtimes.co.uk/jfk-files-how-soviet-lie-that-cia-...
That was my journey as a self-taught programmer who at one point realized had to learn the timeless foundations.
And then MIT's 6.005 [1], where you will apply all this with a realistic language (Java, although the concepts carry to any language), and learn how to design, code and test programs that "have no bugs, are easy to understand and ready for change".
And also learn a bit about algorithms, I don't think that one can design and understand properly without. And what is O(2^n) today will still be that in 100 years. MIT's 6.006 is amazing, both professor and TA [2].
I've seen Knuth's TAOCP recommended around here. Don't even consider that, do a course like 6.006 first, the Everest shouldn't be the first mountain you climb. Likewise favour HtDP over SICP at first, I've been there.
If after all this you are still interested in LLMs, I recommend EdX's "Large Language Models: Application through Production" [3]
[1] https://ocw.mit.edu/courses/6-005-software-construction-spri...
[2] https://ocw.mit.edu/courses/6-006-introduction-to-algorithms...
[3] https://www.edx.org/course/large-language-models-application...
[1] https://stanford-cs324.github.io/winter2022/lectures/data/
There is also the deeper MIT's 6.813 " User Interface Design & Implementation" [4]
[1] https://drive.google.com/file/d/1qblmdoZIJQ4rYGgCPLNusEn2yWi...
[2] https://drive.google.com/file/d/1Cvco6r8a6oV3Sy89RKdIfrZCrGi...
[3] https://stellar.mit.edu/S/course/6/fa18/6.170/materials.html
> The new role requires Data Engineers to be strong across several domains, including data modeling, pipeline development, and software engineering.
> comprehensive guidelines for data modeling, operations, and technical standards for pipeline implementation
> Tables must be normalized (within reason) and rely on as few dependencies as possible. Minerva does the heavy lifting to join across data models.
> When we began the Data Quality initiative, most critical data at Airbnb was composed via SQL and executed via Hive. This approach was unpopular among engineers, as SQL lacked the benefits of functional programming languages (e.g. code reuse, modularity, type safety, etc)
> made the shift to Spark, and aligned on the Scala API as our primary interface. Meanwhile, we ramped investment into a common Spark wrapper to simplify reads/write patterns and integration testing.
> needed to improve was our data pipeline testing. This slowed iteration speed and made it difficult for outsiders to safely modify code. We required that pipelines be built with thorough integration tests
> tooling for executing data quality checks and anomaly detection, and required their use in new pipelines. Anomaly detection in particular has been highly successful in preventing quality issues in our new pipelines.
> important datasets are required to have an SLA for landing times, and pipelines are required to be configured with Pager Duty
> a Spec document that provides layman’s descriptions for metrics and dimensions, table schemas, pipeline diagrams, and describes non-obvious business logic and other assumptions
> a data engineer then builds the datasets and pipelines based on the agreed upon specification
Such a practice is/was actually quite common in (ex)communist countries specially for seasonal work when regular workers weren't enough, like during the cotton harvest in Uzbekistan, and was also organized by the teachers. Shocking for us but culturally acceptable out there.
This is just a modern version, a stopgap for labour shortage, that happened to make it to the papers only because a western company was involved.
This point is very important, and well explained in Stonebraker's paper "What Goes Around Comes Around". What is most interesting is that he is actually talking about half a century old pre-relational IMS IBM databases, but they had exactly the same issue, hence the paper's title. Codd invented the relational model after watching how developers struggled with the very problem you mentioned.
Stonebraker famously quipped that "NoSQL really stands for not-yet-SQL".
He also addresses the impedance matching issue in the "OO databases" section; there is actually a lot more to it, and he gives it all an insider's historical perspective.
Does this include informacion hiding/encapsulation? (to prevent saved objects' internal representation from being exposed).
Traditional databases don't have an encapsulation mechanism AFAIK, which is one of the reasons for impedance mismatch.
This is important because it is a good practice for client code to make no assumptions about the internal representation, accessing data only via the a public interface.
If it happens to be exposed by the database, the clients can use it in their queries. If the internal representation changed later on, such clients would be broken.
Of course, this can be solved by only allowing data access via, say, well designed restful apis (that don't expose internal details), but this would still provide no guarantees.
How about another reason for impedance mismatch, that of storing objects that belong to a class hierarchy?
Start with relational, understand why it is the reference architecture, and from there the tradeoffs involved and what other architectures bring to the table (columnar, streaming, object, in-memory, array, distributed, blockchain, nosql, etc)
To really understand why you should start with relational, read Stonebraker's classic paper: "What goes around comes around": https://people.cs.umass.edu/~yanlei/courses/CS691LL-f06/pape...
It will teach you database evolution history so that you don't end up reinventing the wheel.
Stonebraker's MIT course: https://ocw.mit.edu/courses/electrical-engineering-and-compu...
There are a few lectures of this course in youtube, not by him: https://youtube.com/playlist?list=PLfciLKR3SgqOxCy1TIXXyfTqK...
MIT's distributed systems course also touches on databases: https://pdos.csail.mit.edu/6.824/schedule.html
Course by one of his disciples: https://15721.courses.cs.cmu.edu/spring2020/
Yet another disciple (in edx too): https://m.youtube.com/playlist?list=PLYp4IGUhNFmw8USiYMJvCUj...
The red book: http://www.redbook.io/
For learning SQL really well, including relational algebra, I like this course: http://users.cms.caltech.edu/~donnie/cs121/
My intention wasn't really to bash on Martin Fowler, the problem is much wider than that, and he is certainly not the worst example, just happens to be the OP's subject.
Let me end by quoting MIT software engineering professor Daniel Jackson, I think that he pinpoints the essence of the problem beautifully (in his book "Design by concept", where, BTW, he credits Martin Fowler's book "Analysis Patterns" influence):
In my work as a consultant, I've been involved in discussions about future products and strategic directions, usually under the rubric of "digital transformation". Companies may be keenly aware of what they're trying to achieve (better customer experience, increased customer engagement, differentiation from competitors, etc), but much less certain of how to do it, and especially how to explore new posibilities and get results and feedback quickly. Too often, the options are cast in terms of technology adoption (selecting from the latest shiny new things, whether mobile, cloud, blockchain, machine learning, Internet of Things, etc). These technologies may have great potential, but they are only platforms, and choosing one with the hope that it will transform your business is no more plausible than expecting such an impact from the adoption of a new programming language or web application framework. A better approach is to focus on functionality, which is the source of real value.
Martin Fowler is dangerous in this regard, there isn't a bandwagon he doesn't jump on making it sound like it is the new normal, and, although, as you say, he does mention counterindications, his presentations are still very unbalanced, hyping what is still inmature and extremelly risky. He is an excellent popularizer, I actually enjoy listening to him, but too many buy it uncritically.
I meantioned CQRS, the wider context is microservices. Thank you for presenting so much evidence, you are right, but the thing is, the wider context where it was said, all the hype, is still missing.
Here is one example from a few years ago in youtube [1], at the height of the NoSQL craze this time. He proudly mentioned how his mates at The Guardian adopted it, because "a news article is a natural aggregate", or something like that.
He does indeed add later, in passing, that "some NoSQL databases are immature , we don't have the tools, the experience, the knowledge to work with them well; we've got decades of experience with sql databases".
It turned out that they were burnt at The Guardian, and wrote about their painful migration to a RDBMS and to safety [2].
Now, one can argue, rightly, that this is just anecdotal evidence against NoSQL. The problem is that he used it as evidence for NoSQL. Anecdotal too. The difference is that the former evidence benefits from hindsight, whereas the latter was premature, a space he regularly finds himself in.
[1] https://youtu.be/qI_g07C_Q5I
[2] https://www.theguardian.com/info/2018/nov/30/bye-bye-mongo-h...
If you want to actually learn what they call "clean coding" (aka proper design, these guys are great at creating buzzwords, and great public speakers), the way to go is "Systematic Program Design" (in youtube [1], same material as edx's How To Code), which is based on the Htdp book ("how to design programs"), followed by MIT's OCW 6.005 [2] ("Software Construction").
Both are free. But who knows this? There is no one to hype them, no "movement" behind what are tried and trusted techniques, no flashy CQRS SOLID acronyms.
These two courses, instead, teach you timeless concepts and techniques that will survive all fads.
[1] https://youtube.com/channel/UC7dEjIUwSxSNcW4PqNRQW8w
[2] https://ocw.mit.edu/courses/electrical-engineering-and-compu...
And that problem is where the input can be modelled as a stream of structured events.
This may sound too abstract, but it includes, for instance, pretty much any file. But the value lies more in that it makes it trivial to design a file format that is easy to parse, and process.
For me the value lies in that in the old days techniques were developed to systematically solve programming problems that will always we relevant, in this case processing a stream of events.
It is very well explained in the follwing MIT ocw lecture (it uses Java, but is actually language agnostic):
"Designing stream processors
Stream processing programs; grammars vs. machines; JSP method of program derivation; regular grammars and expressions"
https://ocw.mit.edu/courses/electrical-engineering-and-compu...
To practice it, you could try "Project 1: Multipart data transfer" in the link below:
https://ocw.mit.edu/courses/electrical-engineering-and-compu...
[1] User Interface Design & Implementation (start by having a look at the readings; includes exercises and projects; have a look at the dozens of detailed examples of the latter to see the process that is taught here [9]).
[2] Software Studio (this one sets UI/UX design, at the end, with excellent videos [3,4], in the context of app design, not in isolation, as an extension to conceptual design [5,6], without which it is claimed that you can't get it right, and for which you need abstract data design [7,8]; very good assignments and project)
[1] http://web.mit.edu/6.813/www/sp18/
[2] https://stellar.mit.edu/S/course/6/fa18/6.170/materials.html
[3] https://drive.google.com/file/d/1qblmdoZIJQ4rYGgCPLNusEn2yWi...
[4] https://drive.google.com/file/d/1Cvco6r8a6oV3Sy89RKdIfrZCrGi...
[5] https://drive.google.com/file/d/1lzAbrqk9LNLhGMLbsTuVgAm6p9_...
[6] https://drive.google.com/file/d/1AZ85ui0k2QpCKgfJZXUPSk-jdSi...
[7] https://drive.google.com/file/d/1axvYX7kSWAmxDU451hyCcG-wt88...
[8] https://drive.google.com/file/d/1ET9UkOCE-l0c-FR_2w0TFudCAn8...
[9] https://docs.google.com/document/u/0/d/1KIIOS3W9vF4RT0PdX6cH...
Both very useful to understand what people really think in Spain about all kinds of topics.