235 karma · joined March 11, 2022
A compiler preserves the semantics - if you compile Scheme to C, you get a real Scheme program in C, with full semantics.
A transpiler does not have to guarantee that. It is much easier to transpile Python to JavaScript than it is to truly compile Python to JavaScript. A transpiler can tolerate some amount of leakage between worlds (think 'arguments' as a variable name in Python vs JS).
It really depends on your needs. Sometimes a transpiler is okay, sometimes totally inadequate.
Both highly recommended.
And yes, the reason is auditability but also bare necessity.
Imagine vendor A publishes fresh files every day, but you don't exactly know when. And suppose that the vendor can (and does) republish some of them at any time, with or without notice. You need to watch those files for changes and update your own sources accordingly.
For the collection and dedup, you can construct a tuple of meta-information that will give you a good idea if the file is the same (name, time last modified, file size, etc). Then only if you're still unsure do you download the file and hash it, compare with what you already have, and make a decision. (What if the file name changed slightly and now you must re-categorize the data in the file, or what if a single row was added or removed, etc).
It's also useful when auditing (tamper-proof) or when re-running or debugging calculations that used the data in question. If a user used the data available at 9AM, which was unfortunately incomplete and led to issues, having a full data lineage really helps. You can trivially re-run the computation with the information available at that time to see what happened. Or you can mark some results as stale, etc.
There are different ways to do this and one is keeping track of your content-adressed files in a relational table and give each row a unique identifier, so that you can tell that this particular row in your AAPL OHLCV data comes from this particular file which was fetched at this particular time etc, etc. Pair that ID with a timestamp and when doing computations you can query for "the latest" AAPL data while storing the exact file ID (time is fuzzy, content hash is not). Look into temporal tables and the like.
I was CTO of a FinTech where I built the whole software stack from scratch: the lessons in the book are mostly correct. I say mostly, because as always, there is a lot of "it depends" to take into consideration for your particular project. For example, I chose to not use event-sourcing to avoid the whole state computation issue. A standard append-only audit trail can do the job.
You can't guarantee exactly-once delivery but you can construct effectively-once processing, and that is what you really want.
Store every request and response : absolutely, and not only when consuming APIs, but when collecting any information from the outside world (and, if you can, also log every intermediate transformation step within your perimeter). Content-adressed buckets + a relational table are great for this.
The text also does not mention anything about data lineage. What happens if a vendor updates some data mid-day that you absolutely need to be aware of? You need to be able to account for that, while also re-playing computations that used the old values and get the same result. It's not a particularly hard problem to solve, but it takes some thought.
BTW, here's another quite nice feature of codeBoot: easy sharing! The following link will bring you to a suspended execution of 'print("hello, world!")' that you can then single-step.
https://app.codeboot.org/6.0.0/?init=.faGVsbG9fd29ybGQucHk=~...
Now I'm very curious about integrating OCaml too :)
Our tech stack is different but the choices we made are quite similar. Multi-tiered platform, markdown for authoring, executable exercises, teacher platform to produce, give, receive and grade homework/exams, etc. A distinguishing characteristic is that codeBoot's Python interpreter (pyinterp) allows single-stepping through the code. That's quite useful for teaching and studying.
We have a few exciting features coming up and we're working on a proper landing page and clean English translation for the book. If anybody is interested to learn more, reply here or contact me (email in my profile). I'd love to connect with educators, students or hackers alike!
I was CTO/engineer at a FinTech up until last month. You can find my email in my profile.
Remote: Yes
Willing to relocate: No unless very strong commitment.
Technologies: Python, JavaScript, Scheme, PostgreSQL, ClickHouse, compilers, interpreters, foreign function interfaces, FastAPI, Flask, Celery, RabbitMQ, Redis, Pydantic, SQLModel, technical strategy, team management (6+ devs + interns), SDLC best practices, etc
Résumé/CV: Sent by email
Email: See profile!
I have over a decade of experience working in various roles and industries, from healthcare to education to finance. My last role was CTO of a FinTech that I brought from zero to one, architecting and implementing all of the tech stack from compliance to distributed locks while managing a team of 6 developers.
I am looking for interesting problems and a good team. I strongly believe that good, thoughtful work pays dividends when things break or need to move fast. Boring, proven tech before shiny new thing.
And to clarify: I didn't "forget" how to code. I froze for a moment, after relying on a tool to write for me for most of a year. More like rusting.
Except not everything was properly documented, and it turned out the employee had given admin rights on some resources to a contractor which proceeded to wreak havoc on their behalf (the 'rm -rf' kind). Eh!
Would you say calculators rot brains?
Did the invention of writing rot the brain?
The skill is in making the LLMs reliably generate useful and pertinent streams of tokens. That takes work, reading the output, intuition, experience, rigor, real commitment to doing good work, not fall prey to being lazy, etc.
Just yesterday I was interviewing for a very interesting job and I completely flunked the coding question in an unacceptable way for my level of experience. The question was easy, I just couldn't get past some syntactic issues. For 8 months, Claude wrote all of my Python classes and Pydantic types. Now I had to write a dataclass, and because I always just resorted to standard classes before the advent of LLMs, I stumbled. And froze. And panicked. And that was it. Of course you could say I should have just scrapped the dataclass and written it as a simple class. The point is I felt very, very stupid. LLMs suddenly felt like a huge disadvantage.
All this to say I disagree with LLMs "rotting" my brain. Quite the opposite, I know that it's possible to use LLMs to be efficient and correct. It's more the actual mechanical act of writing that gets rusty.
Remote: Yes
Willing to relocate: No unless very strong client commitment.
Technologies: Python, JavaScript, Scheme, PostgreSQL, ClickHouse, compilers, interpreters, foreign function interfaces, FastAPI, Flask, Celery, RabbitMQ, Redis, Pydantic, SQLModel, technical strategy, team management (6+ devs + interns), SDLC best practices, etc
Résumé/CV: Sent by email
Email: See profile!
I have about 14 years of experience working in various roles and industries, from healthcare to education to finance. My last role was CTO of a FinTech that I brought from zero to one, architecting and implementing all of the tech stack from compliance down to distributed locking while managing a team of 6 developers.
I am looking for interesting problems and a good team. I strongly believe that good, thoughtful work pays dividends when things break or need to move fast. Boring, proven tech before shiny new thing. Sleeping at night is priceless.
Contact me, let's talk :)
We don't want to break the current experience but we need to do a better job of explaining what our software can do.
I appreciate your comments!
It really is a joy to program with, but we're struggling a bit to communicate everything it can do. We are working hard on that front and should have a landing page and better explanatory material soon. We're very interested in feedback. If anybody wants to learn more, just contact me through the email in my profile.
Cheers!
It's a fully client-side Python IDE with single-stepping, a virtual (non-hierarchical) filesystem, an FFI to call JS code and a few other things (see the docs). Sharing apps in CodeBoot is trivial: right-click the "play" button and copy a shareable URL. I have helped people solve data wrangling problems using CodeBoot and they now have their little app bookmarked. It works really well.
I could go on for a while. AMA if you're interested. We're actively working on it and some great new features are on the way!
Pr. Marc Feeley's lab develops codeBoot [2], an online IDE to teach students programming (and more!). We created BLINX as a hardware platform for students to go along with our IDE. The device acts as a data collector for various Grove sensors and publishes the data as an HTTP endpoint. You can program it directly from codeBoot.
BTW if anybody has any questions feel free to reach out!
[1]: https://www.linkedin.com/company/blinxinc (working on a landing page)
[2]: https://codeboot.org (also working on a landing page)
I understand your comment was tongue-in-cheek, but we certainly have an interest in cross-language interoperability! You can check out our work here:
- https://try.gambitscheme.org is Gambit compiled to JavaScript with the universal backend. Evaluate \alert("hello!") at the REPL to see the JS<->Scheme Syntactic FFI in action.
- https://codeboot.org is our own Python interpreter running in the browser. It has a Python<->JS FFI. Evaluate \alert("hello!") at the REPL to test it out. You can even import JS libraries using the standard Python syntax by replacing the identifier with a string: import "https://mycdn.com/mylibrary.js".
- https://github.com/gambit/python is a Gambit module that integrates Gambit with CPython, using the same syntactic FFI. You can import PyPI modules from Gambit.
References to conferences/papers describing these features can be found on my GH profile (https://github.com/belmarca). AMA if you wish!
Remote: Yes
Willing to relocate: No
Technologies: Python, Scheme, JavaScript, C, Vagrant, Docker, Ansible, PostgreSQL, VueJS, AWS, OpenBSD, ESP32, DICOM, Healthcare, Compilers, Interpreters, Dynamic languages, Web, Full-stack, Backend, Web3, Blockchain, etc...
Resume: On request
Email: marc-andre (\dot) belanger (\at) umontreal (\dot) ca
The list of keywords is not in any particular order of importance or experience. My most recent (and current) experience is working on a full-stack Python/JS code base where we develop a web-based Python interpreter. Serious inquiries only please.