174 karma · joined March 18, 2017
In 1799, Paolo Ruffini published a 500 pages long proof showing that there is no closed algebraic solution for the roots of a polynomial of degree five or higher. The proof is extremely verbose and brute-force, essentially enumerating and checking hundreds of cases by hand. It is by today’s standards insignificant.
About 25 years later, Evariste Galois proved the same result in about 95% less space by describing the first general theory of groups and fields. It is considered one of the greatest contributions to mathematics of that century, not because of the result, but because its approach opened up a whole new universe of questions, methods and insight. There would be no AES encryption without Galois.
To me, Astras proof looks like Ruffinis proof.
> I'd settle for a good way to express said tree in a plain text file
What you are referring to is commonly called an IR (intermediate representation). Compilers typically take the AST generated by the parser and translate (“lower”) it into increasingly hardware-near and optimized IRs before performing instruction selection.
I briefly browsed through the links you provided. IIUC, BitGrid is essentially one of the Turing-complete 2-D cellular automata. This Reddit thread might be interesting to you [0].
Regarding transactionality: There is an entire area of research on "hybrid transactional and analytical processing" (HTAP) systems that unifies OLAP and OLTP systems. Hyper [1] pioneered this path at TU Munich, it's successor Umbra [2] recently incorporated as CedarDB [3]. There are lots of others. Most of these systems, AFAIK, are relational.
Regarding data model: What we've seen in the past few decades is that non-relational DBMS (excluding key-value stores) only make sense in rare edge cases that require huge scale. There has, e.g. been research [4] that shows that graph databases are still, well, lacking, compared to relational systems. The common pattern seems to be: unless you need to service very specific workloads at huge scales, SQL is probably enough [5]. Then again, it really comes down to intrinsics. If you were to, for example, implement distributed locking using Postgres, you would likely run into problems with MVCC and Xids very quickly.
So, as you already mentioned, there is no silver bullet. But even today, unless you are Meta or Google, SQL is probably enough for a long time and lots of use cases.
(Full disclosure: I'm working on Hyper full-time).
[1]: https://hyper-db.de/ [2]: https://umbra-db.com/ [3]: https://cedardb.com/ [4]: https://homepages.cwi.nl/~boncz/edbt2022.pdf [5]: https://www.youtube.com/watch?v=VxKt245X_ws
And Hyper is alive and well at Salesforce/Tableau! The team working on it is still in large parts the original Hyper team from TUM. You can actually download Hyper (as a binary with language bindings) and play around with it [2] for non-commercial use cases.
If you think Hyper/Umbra is cool, the TUM database group has lots of other very interesting projects going on at the moment. LingoDB [3] pushes the database-as-a-compiler idea to the extreme by implementing query optimization and compilation query compilation in MLIR. LingoDB is open-source. Also Viktor Leis, who stands behind (among many other things) Hyper's Morsel scheduling and ART indexes as well as Umbra's buffer management recently started a very interesting project [4] to heavily co-design the DBMS together with the OS in a unikernel approach. Really interesting stuff!
Disclaimer: I work on Hyper. Views are my own.
[1]: https://cedardb.com/ [2]: https://tableau.github.io/hyper-db/docs/ [3]: https://www.lingo-db.com/ [4]: https://www.cs.cit.tum.de/dis/research/cumulus/
Any insights on the business model? Is this just to advertise the pro tier or do they monetize usage in another way (e.g. selling training data)?
However, the crux is in the details:
> You can increase the % enough so that overall demand for developers goes down or doesn't grow as much as it would have otherwise.
I would be at least skeptical of this. Every push for commodification that we've seen in the software space so far has been absorbed by demand. Will this continue forever? Nobody knows. At least where I work the backlog is filled to the brim, and every new iteration of tooling begets more babysitting to unlock the promised gains. And customers still have a never-ending list of hyper-specific feature requests.
The friends and colleagues at the Senior/Staff level who are using Copilot/GPT-4 (and have admittedly become much better than me at prompting) didn't exactly become "hyper-productive". Sure, they get code pushed out faster, but they still work long hours and complain about deadlines.
This is not to say that we're all fine forever and things will not change. But as long as we don't experience an across-the-board temperature shift in the job market decoupled from macro-economic events I wouldn't put too much attention there. In the end, doom scrolling is also just a form of procrastination.
Look — I appreciate that you want to help OP out here, but please keep HN free from low-effort LLM-generated answers like this. Everyone here has access to chatGPT, so the added utility of pasting responses from it is close to zero.
[1] https://code.visualstudio.com/docs/python/jupyter-support-py...
To rephrase my comment above: I don't want to blame the team behind Copilot for not getting everything right on the first try. Neither am I in a position to do so, nor would I want to live in a world where smart people aren't allowed to make mistakes.
What irritates me is that there are two possible scenarios here:
1) They knew about potential issues and decided to release it anyway (without at least addressing them verbally). 2) They didn't.
And frankly, I don't know which one I like less. Even though it's still a beta/preview, either option seems to signal a degree of negligence? that feels unnerving given the potential impact of such a system.
That being said, if we do live in scenario 1) than I am certain that better framing could have prevented the PR fallout that we're seeing right now (at least partially). IMHO, GitHub (the platform) is still a great product after all.
Don't get me wrong, the folks behind Copilot are clearly, without any doubt smart, creative, and capable. But then... None of these issues (reproducing licensed code ad verbatim, non-compiling code, getting semantics wrong, and now this) are 0.01% edge cases that take specialized knowledge to see or trigger. I remember some of them being called days ago in the initial HN thread by people who haven't had beta access.
I really wonder how this announcement/rollout looked like on the management side of things. Because a) these shortcomings must have been known beforehand and b) backlash from people who feel threatened for their jobs/"stolen" of their open source work was (I guess) foreseeable? I've already read calls to abandon GitHub for competitors; this can hardly have been an acceptable outcome here.
Nevertheless, Copilot is still one of the most innovative and interesting products I've seen in a while.
I met a lot of brilliant people from all over the world during my undergrad at TU Munich; many of them stayed. Google, Apple, Microsoft have expanded rather aggressively over the past years, and the startup scene is growing. Berlin, Hamburg, Frankfurt seem to follow suit.
Granted, it's much less likely (and also less politically incentivized) to become a "millionaire employee", but I guess ultimately, the correlation between wealth and wellbeing is simply weaker in (northern) European economies? As long as immigration policies don't change (and why would they? the current model aligns well with lobby incentives), I don't see much reason for this to change.
Also, Poland, Romania, Bulgaria, Spain, and Greece have a different political and economic history than Germany, France, Denmark, Sweden, etc.
The workflow for creating a new project looks like this:
1. Create a project directory (e.g. 'myproject') and `cd` into it. 2. `git init` 3. Fixate the Python version for that project with the `pyenv local` command (e.g. `pyenv local 3.8.6`). This creates a `.python-version` file that you can put under source control. Within the `myproject` directory tree, `python` will now be automatically resolved to the specified version. Your system Python (in fact, any other Python versions you might have installed) remain untouched. 4. Create a new poetry project (`poetry init`). This creates a `pyproject.toml` which contains project metadata + dependencies and can also be checked into git. 5. Add dependencies with `poetry add`. Here, you could for instance add Jupyter Lab (`poetry add jupyterlab`).
To access installed dependencies, such as the `jupyter lab` command, you can either execute one command in the virtualenv directly (`poetry run jupyter lab`) or spawn a shell (`poetry shell`). If you open a Jupyter Notebook that way, the packages installed in the virtualenv are directly available from within Jupyter Notebooks, without having to mess around with installing IPython kernels.
I like this approach, because it gives you full flexibility, while being portable and easy to use. It gets you around having to deal with conda (which I found to be frustrating at times). Also, you're not tied to the Jupyter frontends, but could e.g. just install `ipykernel` and open notebooks in VSCode.
[1](https://github.com/pyenv/pyenv/) [2](https://python-poetry.org/)
Edit: Moved the links
The strategy agreed upon today was a revision of an earlier plan that would have cost a "mere" 5bn euro. This deal was however opposed by the individual federal states of Germany, which demanded 60bn – just to put the final 40bn in perspective.
40bn is still a huge sum of money (about 1/3 Jeff Bezos, for scale), but as a German voter and tax payer I have to say that I am entirely fine with this. Germany is a very prosperous country, and we have to get away from coal, the sooner the better. True, it probably could have been done with less money, but to me 40bn is still a favorable deal over no deal at all.
What helped me (as with many proofs and concepts in Math) was an image, a visual metaphor if you like.
Imagine a very, very large paper on which you place infinitely many dots in a grid. That's the infinity you were referring to, the infinity of a for-loop, discrete infinity. Here's the trick: you can always add more dots, say by making the distance between grid points half as small, which would quadruple the number of dots in your grid. But no matter how many dots you place on the paper, ho matter how fine your grid, there will always be holes (imagine "zooming in" on a square of four grid points). In fact, most of the paper will be empty!
The other kind of infinity, continuous infinity, does not have any holes. Every spot is covered. You could not add any grid point, because the whole paper itself is painted.
I'm not a "full-time Mathematician", so this view may be entirely wrong. But it helped me understand and appreciate Cantor. Perhaps it did the same for you.
Cheers!
AI is academic (as a synonym for 'theoretical' and 'math-intensive'). Once you look beyond purely symbolic AI, which proved to be infeasible as @curuinor pointed out somewhere here, you will need to build up at least basic knowledge in probability theory and linear algebra.
The path I'm following at the moment is a quite rigorous one and is outlined here (http://www.deeplearningweekly.com/pages/open_source_deep_lea...).
If you've never had any exposure to probability theory or statistics, I recommend having a look at the course "MIT 6.041 Probabilistic Systems Analysis and Applied Probability" taught by John Tsitsiklis at MIT (video lectures are available through YouTube and MIT OpenCourseWare for free). Both the course and Tsitsiklis' book are superb learning materials to get into probabilisitc thinking.
Edit: Link was broken. Thanks to @blauditore.
If you're also looking for a course that goes alongside the book, I highly recommend UC Berkley's CS188 (you can find it at http://ai.berkeley.edu).
The lecturer Pieter Abbeel does such a good job explaining stuff and the programming exercises are really neat.
Edit: Formatting