Take more screenshots
alexwlchan.net
alexwlchan.net
It's amazing how many emotions seeing that one image gave me. But the biggest was just this overwhelming sense of nostalgia. As I looked at that, I could remember what I was thinking, what I was feeling, everything that was happening in my super confusing teenage life at that time. Occasionally I will look at that image now, even 22 years later, I can still feel all those feeling again.
Of course, my ex's character is in the screenshot too. So, a bit bittersweet as well. :/
[1] https://en.wikipedia.org/wiki/Madeleine_(cake)
[2] Irma Rombauer, Marion Rombauer Becker, and Ethan Becker. Joy of Cooking (rev. ed.). Scribner. 1997. pp 962-963.
I have deleted everything from uni though.
Especially when you see parts of a conversation with someone you didn't talk to in over a decade, or who passed away. You'd think that's what a photo would do, so I was surprised how strong of an emotional reaction screenshots can trigger. But then if you were spending 90% of your time online as a teen/early 20s it's not that surprising on a second thought.
You know, this is a very enlightening point. I never really thought about it from this angle, but there is a lot of truth to this.
I struggled a lot as a teen with anxiety, depression and bullying. I had a few IRL friends, but the very vast majority of my social interaction during that time came via MUDs and chatting. Many of the people I played and chatted with were fellow social outcasts, and we created our own parallel virtual communities to support and lift each other up. It didn't matter where we were, what we looked like, or how we did or didn't fit in. Many days in the 90s it felt like going to school was the thing I had to put up with, and logging in and seeing my friends when I got home was my real life.
Without them, there's a very real chance I might not be here today. Even all these years later, the people I met virtually during that time are still some of my best and closest friends, and it's a real treat when my travels take me close enough that we can meet for coffee or lunch. Many were at my wedding even, and in one case that was the first time I had ever met them IRL. And yet we knew each other deeply. It felt like we all grew up together because, kinda, we did.
When you look at it like that, those of us who grew up in that environment would look at a screenshot from that era the same way others might look at random photos from high school. Because this was our world.
* It's a mental hack to keep me accountable, especially now working from home. If I'm in an office anyone can look over and see whether or not I'm working. It started as an attempt to mimic this feeling at home, even though I'll be the only one to ever see the recordings.
* It allows me to go back and see how I worked in the past. I have a few videos of myself working from 2015 which I think is pretty neat just because of how different my workflow was back then compared to now. I'm not using the same tools or even on the same operating system.
* I'm working on video games which is what makes this very useful for me. If something visually interesting happens, or if there's graphical bug of some kind, I can go back and breakdown exactly what happened. I've stepped through videos frame by frame in the past to debug, it's been surprisingly helpful.
* It allows me to go back and see my progress. I can know what I was working on a given day, see how far I've progressed, it's just generally a good motivator. You can of course do this with git, but if you're working on something visual it can be nice to see it in motion rather than a textual diff.
I think I’d prefer live streaming to keep myself accountable to one-on-one sessions with both cams on.
But the TLDR, ffmpeg can remove duplicate frames:
ffmpeg -i in.mkv -map 0:v -vf mpdecimate,setpts=N/FRAME_RATE/TB out.mp4You can go even lower with other encoders (x265) + if you don't record audio at all
That's only 1.14 TiB per year doing it 5 h/week * 5 days/week * 52 weeks.
- Of the lossless encoders in OBS/libav, utvideo was the best in both CPU usage and efficiency, followed by lossless ultrafast x264.
- An SSD can handle even uncompressed 24bpp 1080p60, which is 375 MB/s. Typical screen content compresses well below the ~100 MB/s write speed of an HDD. Fullscreen video does not, instead gradually filling the write cache until either OOM or thrashing.
- For onscreen content, I prefer the bitrate tradeoff of keeping PC color range and not chroma subsampling.
This technique isn't as effective for this use case of recording several hours daily, since reencoding must be fast enough on average to keep up. Best to already have a home server (any spare desktop). Otherwise, use AOM codecs, known for poor multithreading, to encode at full speed without hogging CPU.
ps, temporal compression means that dropping framerate makes surprisingly little difference with modern codecs. But I really should be writing down the results of my ad hoc tests...
Well, it can for an hour.
Yea, on second thought, long term screen recording like this would be a terrible waste of silicon. At 150 MB/s, 46 days of nonstop recording would exhaust the entire 600 TBW of a 1 TB consumer SSD. And for fun: a high-end 500 GB SSD writing at 3500 MB/s could burn its 300 TBW in a single day, not counting the EoL slowdown.
Curiously, RAID 0 hard disks are perfect for such workloads, yet the blogger has still chosen to use consumer grade SSDs.
I do this sometimes. the accountability hack works even better if someone could be watching.
Bonus points is that it feels way more natural to narrate aka rubber duck problems when you are streaming.
Screen captures also compress much better than live action since most frames are duplicates of their predecessor. So a cheap NAS can run for years before you start thinking of deleting VODs.
Why? I REALLY enjoy the dopamine rush when you are struggling then find a solution. I see myself pulling my hair, staring blankly at the screenshot then at a random moment of pure luck I find a solution and it literally is euphoric.
I enjoy relieving those moments.
Game development where GPU is not really going spill the beans of what is happening under the hood that easily - one can stuck for a longer times easily.
I've been working on this problem for years and I've gone through dozens of design revisions. Sometimes I catch bugs in the design phase. Sometimes I only notice core problems after I finish implementing parts then attach a fuzzer. (Fuzzy boi is the best and the worst.)
Some design problems have taken me months to find a solution. One of my solutions ended up with me rewriting of thousands of lines of working code, written over several months.
We're making great progress; but sync engines are hard.
I used to think that video formats would no longer be supported over time, but even the oldest weird video formats still play in VLC and MPC, and probably would work fine if uploaded on YouTube.
During the stream, I keep up a fairly constant spoken description of what I'm doing, what I'm thinking, what problem I'm stuck on, etc.
I've noticed I've also been speaking my thoughts out loud when programming, but not streaming. It ends up being a continuous "rubber duck" conversation, and feels (completely subjectively) like it helps me develop easier/better.
I'll try it!
I've used this before on engineering grade machines, but it doesn't do so well on "everything is in the cloud so you can use a word processor quality" laptop, any advice?
I have a small shellscript that takes all video files recorded by OBS, and runs them through this ffmpeg command:
ffmpeg -i in.mkv -map 0:v -vf mpdecimate,setpts=N/FRAME_RATE/TB out.mp4
Using mpdecimate removes duplicate frames, so if nothing is happening on your screen (although smaller changes gets ignored, like my clock showing the seconds), it removes those duplicate frames.So one ~1 minute video of you thinking for 40 seconds can get reduced to 20 seconds. Not uncommon for some of my video files to go from multi-GB to just ~100 MB when removing all the pauses.
What about hard disk space? How you handle it?
"ffmpeg -f gdigrab -i desktop -c:v libx264 -preset medium -fps_mode vfr -crf 0 -an -vf mpdecimate capture.ts" (Windows, -f x11grab -i $DISPLAY for Linux on X11) produces a lossless video that averages 102 MiB/hour for me (1920x1080 @ <= 60fps). That's about two cents a day at current disk prices, and easy to upload as a private YouTube video if you don't want to lug the files around.
If you don't like the CPU hit, use -preset ultrafast to record the capture using less processing power but giving a larger file, and then re-encode that file later using -preset slower. There's no quality loss if you used lossless mode (-crf 0), and for content like this the savings are especially large (reducing to around one-quarter ultrafast size, in my experience).
Legally it’s very helpful if you create a paper trail of your work
I never really regret deleting something, but that could be because I try to keep my life simple, within reason, and focus on the future.
I also recognize, as you point out, that there are limits to this -- sometimes there is a genuine need to keep a record. As a programmer, my work is all tracked in git. For a creative professional, I assume that a basic requirement of that sort of job is an excellent backup system.
I learned that the nostalgia is not about the files by itself but my life context at that time. I don't miss old code or Old OS's. I miss that sense of wonder when I was less experienced and more naive, and everything was new.
I'm learning the same - whenever I feel nostalgic about playing an old SNES or PSX game, I've realized that it was just about that time in my life, and usually just watching a clip on youtube or listening to the soundtrack is enough to scratch the itch, rather than actually playing the game again
I find that horrifying. I would probably scramble to delete that as soon as possible. Don't know about anyone else here, but having my junk float around the internet is mortifying. I delete unused accounts as soon as the thought of it pops into my mind.
1. Photo apps like iCloud Photos, Google Photos, and Photo Prism are getting much better at auto organizing/cataloging what photos are and surfacing them together in much more interesting ways. More photos is now a plus instead of a minus. 2. You can always delete, but if you never record you can't go back and record (often).
So I take a Marie Kondo "Keep what Sparks Joy" approach and delete anything I find that I do not care for when I find it. I also sometimes pick an area of stuff I have and try to aggressively delete things I don't care for.
In day to day life these past record could as well not exist, but you still get to go look at them if you're willing to make the effort to do so. Also being willing to lose these items if something catastrophic would happen relief a lot of the archiving burden.
I'm sure your parents or someone from their generation have actual, physical photo albums from their past - and that the experience of browsing through these photos brings back things that they haven't necessarily _forgotten_ about, but that they wouldn't have brought into active memory unless they were browsing through them.
Over the years I have experienced and built many things that I do not deem "important" to me, yet when I see them mentioned (even in writings by myself), it takes me back to that point in time - all the feelings, learning and discoveries that it brought to life.
An example is scrolling through a list of my repositories on GitHub. Some of the projects on there I have "forgotten" about, but with the mention of it I am instantly brought back and remember a whole lot more details - motivations, feelings, the ecosystem...
I.e. photos and important documents are only about 50GB so very easy to keep. This would be the things that I wanted my family to keep, If I died.
Then theres random projects and files that I keep in yearly archives, just to look and remember what I was up to x years ago. That's also not huge amounts of data (1-2 GB per year).
And then theres data which I didn't create. E.g. movies, game installer,... Where a loss wouldn't really be a big problem.
Basically, if it doesn't need to be specifically filed somewhere, I just put it in my "downloads", knowing that it will be somewhat searchable in the archive by either date, file type or in-document search. This is great for all of the bits and pieces that you don't necessarily have a home for or want to manage.
I'd like to have the names of all programs visible in the screenshot (easy), possibly application specific metadata like the opened filename or a URL (more difficult) and more generally full OCR of the visible text (pretty easy). You'd need a PDF to get the most out of this, but presumably most other image formats have generic metadata storage.
Regarding open programs: I once led a project where we developed something that keeps track of the software you're running (no screen recording) in order to conduct research into attention and distraction. We didn't have the resources to support many platform versions, so we wrote only a Windows client (most used OS on the floor). It was similar to RescueTime https://www.rescuetime.com/ but more respecting one's privacy and absolutely avoiding the cloud, as we deployed the experiment in a lawyer-intensive environment; for instance we logged running program names but not titles of open windows because file names often reveal sensitive matter. We questions occasionally how productive they felt in the last half hour, and they could comment.
Why can't we le the ML models figure classification themselves and then give them the human data to adjust to be readable by the human.
You can set the File name using various parameters in combination like - time date - program name - window title
This can also be configured to do complex workflow like
- for each screenshot add a border then copy image to clipboard and then upload to Imgur and copy url finally run and ocr and copy results as well.
These archives saved me from data loss a few times (due to powerloss, or bad mistakes). I just browse through the archive to look for the screenshot where I was coding to recover them by typing it again.
this is the screenshot tool I use https://github.com/soruly/TimeSnap
I would also like this, but what worries me is privacy. There’s already so little that it makes me wonder whether this is degrade it even faster.
Consider that with a text editor like Vim, for example, you can "time travel" [0] through your file's edits, or even have undo branches/trees [1][2] available per file. That saves you the trouble of having to transcribe text from screenshots, and also barely uses any storage space.
Plain text is also highly more portable and more likely to be recoverable in case of drive failure or file corruption.
Additionally, or alternatively, you could try any sort of manual versioning system or background automatic backup solution that keeps versions of files as you work on them.
[0]: https://vimtricks.com/p/vimtrick-time-travel-in-vim/
When we upgraded to System 7, the version of MacPaint didn’t run and he told me that essentially the art was lost and unrecoverable. I can picture in my mind what they look like, and wish I had the file to look at.
But then I realized most of the "traveling" that I do is the different places that I go on my computer. If I'm going to take photos when I travel, I should also be taking screenshots constantly, to document places I went on my computer. Such screenshots are useful because they will capture:
1. the large number of websites that exist in a particular year, but which won't exist 5 years later
2. minor interests I have one year that I don't have later. Occasionally I'm curious when I first got into a particular interest, and seeing the screenshots is helpful for documenting that.
When I go back to the oldest entries of my weblog, from 2005, I notice that more than 80% of the links are now 404. The Web is constantly disappearing. Like a forest on its way to extinction, you might as well photograph it now, because it won't be there 10 years later. Most of the websites you visit, you cannot go back and visit them a few years later. Take a photo of them while they still exist.
That's just my opinion, obviously. But I find the obsession with nostalgia in our culture to be sad and destructive.
My occassional evenings spend looking back through the life that brought me where I am are sad and destructive?
I find them to be very effective tools for reflection.
I delete everything when I'm done with it. I backup less than 2 GB of personal data. I have around 100 photos of my entire life, and I'm certain I'll get around to deleting all of them too.
I could probably create something with cron, but maybe a neater solution already exists.
Source:https://github.com/wanderingstan/Lifeslice
Download page: http://wanderingstan.github.io/Lifeslice/
I developed it as an early Quantified Self tool primarily for the Webcam shots, but also have been saved on more than one occasion by having screenshots of work that would otherwise be lost.
Edit: the first version was just a shell script, which if you want a starting point to modify: https://github.com/wanderingstan/Lifeslice/blob/master/1.0-S...
while sleep 5
do
import -id root `date -Is`.png
done
Then make a time lapse with mplayer et.al. $ which import
import not found
How should I install that tool?/$INTERVAL * * * /usr/sbin/screencapture $CRONSHOTSDIR/`date +\%s`.png
and then to make the timelapse:
ffmpeg -r $FPS -pattern_type glob -i "*.png" -vcodec libx264 lapse-`date +%s`.mp4
You have to give screen capture permissions to cron.
On macOS here's a little shell command to take a screenshot every 60 seconds and place it in the /tmp folder:
while true; do
screencapture /tmp/screen_capture_$(date +\%s).jpg ;
echo "wrote screenshot";
sleep 60;
done
You can also include the name of the front most application by fetching it with AppleScript and then putting that in the filename. (Or you could put different app screenshots in different folders) while true; do
FRONT_APP=$(osascript -e 'tell application "System Events" to name of first application process whose frontmost is true');
screencapture /tmp/screen_capture_${FRONT_APP}_$(date +\%s).jpg ;
echo "wrote screenshot";
sleep 60;
done
Of course if you want this to run long term, use a cron job or even better, a launchd job. With launchd you get better coalescing and sleep recovery behavior. You can also add other fun trigger events. Like take a screenshot every time your `.zshrc` is modified [0]. Or a screenshot every time your internet access changes (by triggering on /etc/resolv.conf)[0] Look up `WatchPaths` on my favorite launchd guide here: https://launchd.info/
Unfortunately i only see the old version here with flameshot taking screenshot at full resolution.. my few later versions turn screenshot to black and white and applied a few imagemagick tweaks to make screenshot file incredibly smaller to store but you get the idea :): https://gist.github.com/santrancisco/9d14e0105316cfa15f98f0f...
I just went through three years worth of screenshots from attending tech bootcamp and working my first dev job. It was a great reminder of projects and people I care about who I'll probably never get to work with again. Like random office polaroids for the remote work era (as if I'm that old).
To my own surprise my only regret was that I should have taken more screenshots... Also they're all pngs...
Overall though it's nice to see old screenshots even on my phone, but intentionally "taking a screenshot for the memory" feels weird as every one of my screens is boring on its own.
I consider this an adjunct to my photography hobby. When I travel, I take opportunistic photos of interesting things I see. When I use the computer, I am a traveller in a digital universe and I do the same.
I find it really interesting to see the evolution of desktop UIs that I have used over the years, and to see the slow change in applications I typically use.
I kept literally nothing from that time or probably even a decade after.
Literally hundreds of thousands of lines of code.
The bigger thing to emphasize I think though regardless of archival method is to make sure to regularly back things up. Some of my earliest things from the 90s were on the boot drive and got wiped when the family computer needed a reformat. Later on when I had my own computer, a lot of stuff was on an external drive to make room on the boot drive, but one day the external decided to kick the bucket and everything on it went up in flames because there were no other copies. I didn't have much cash at that point since I was a high schooler but I'm sure I could've figured out something that would've preserved at least the most prized documents.
These days I have everything automatically incrementally backed up with Backblaze but now that Time Machine on macOS uses APFS snapshots and is more storage efficient I also want to use my home server for backup.
Most screen recorders are incredibly cludgy. They either require extra cropping and editing after the fact or tank framerate into being unusable. I don't understand the technical problems and the whys though.
Also, it packages pictures into a PDF which I find to be more organized than a bunch of pictures scattered about. I usually close the PDF every day and so each PDF corresponds to pictures taken during a day which gives it a context to make stuff easier to find.
https://www.virustotal.com/gui/url/718fe790bb2c1d97428199830...
tl/dr -- virus total is happy with the file.
It remains possible that middleware of some sort has altered the file before you accessed it.
You can check the hash of the file like this:
get-filehash ".\TimeSnapperProSetup.exe"
and it should return "B50A4449C9C36871280A842A530EA19694C7CBC4CF55CED5A620D8838D87CA1E" -- the same hash shown in the virus total report.
It feels like a superpower sometimes... Reviewing the exact research steps of projects from years ago based on a timestamp from browser history, digging up an archaic screenshot of just the right configuration screen based on a file modification date, etc.
Sorry you find it useless...
I managed to get most of the data off my old Apple IIGS hard drive but there were definitely block errors. So I can still boot up the desktop I had in the 80's and 90's in an emulator. Brings back so many memories. http://www.oldcomputerstuff.com/resurrecting-my-apple-iigs-a...
Plus just the feeling of WIP app I guess. That could have been me.
Being digital is what makes the idea of an "exact copy" make sense. You can make an exact copy of some version of the Torah or the Symposium because it's only the discrete letters that matter; the analog nuances of tone of voice or thickness of pen stroke do not count.
So digitality is the alternative to the ephemerality of the analog, which is inevitably eaten up by moths and rust. We all know this about digitized language, but for some reason now that we've digitized reasoning in the form of computer programs, we habitually throw up our hands and declare defeat in the face of inevitable ephemerality.
This is bullshit.
What I really want, instead of screenshots, is a deterministic, reproducible computing environment. The idea is something like uxn or Nock: a platform that's simple enough to stay compatible forever, and efficient enough to be used for many things, even if there are a few things that I do on a computer that need more performance.
There are a lot of inspirational examples that offer tempting evidence that this is possible for large, interesting classes of computations: the Smalltalk-78 revival emulator Vanessa Freudenberg wrote, the UM of the Cult of the Bound Variable (which had over 300 successful independent reimplementations), Nguyen and Kay's sketch of Chifir, Lorie's archival UVC, Wirth's RISC, uxn/Varvara, the JVM, and the numerous emulators of things like the MS-DOS environment, the NES, and the Gameboy that are good enough to run the original games.
I'm not saying it would be an improvement to do all your digital creative work on an emulated Gameboy in order to ensure that it was reproducible. I think we can do a lot better than that. None of the presently existing archival virtual machines are adequate. But I think the reproducibility of Gameboy games tells us that we don't have to accept bitrot as the price of using computers.
Alex says, "They’re not as good as having the original, working thing – but they’re much better than nothing". Well, let's figure out how we can have the original, working thing! This is software, it's a simple matter of programming.
Human genome is, effectively, a historical digital record (Nature, DOI: 10.1038/nature10231).
Many aspects of life, however, are analog. Magnesium concentrations, membrane polarizations, molecule orientations, temperature, and so on.
There at least three small Wasm interpreters in Rust
4kloc https://github.com/yblein/rust-wasm
3kloc https://github.com/k-nasa/wai
500loc https://github.com/rustwasm/wasm-bindgen/tree/HEAD/crates/wa...
This list has 169 Wasm instructions https://github.com/rolfrm/wasm-lisp/blob/master/instruction....
Wirth's RISC is neat, I'd love to re-do it in RISC-V (only 47 instructions in the base ISA). UM, Chifir and UXN look like Art (not pejorative), I'll definitely read the Chifir paper. They would be great systems to run on top of Wasm.
https://git.sr.ht/~bctnry/chifir
One might be able squeeze a Chifir VM into an ESP-32 (with external PSRAM).
And I appreciate you setting me straight about the extent of wasm's instruction proliferation.
But I find much to disagree with.
— ⁂ —
> There at least three small Wasm interpreters in Rust
None of those are small; the smallest one you found is 500 lines of code, and it's very incomplete, implementing only 11 of the 436 or however many wasm instructions there are, and even those it only implements partially. (Its only arithmetic is addition and subtraction, for example.) Even the 3kloc one says it doesn't pass the wasm testsuite; its pub enum Instruction has 174 items and most of those are implemented as follows:
Instruction::F64Store(_, _) => todo!(),
Instruction::I32Store8(_, _) => todo!(),
Instruction::I32Store16(_, _) => todo!(),
Instruction::I64Store8(_, _) => todo!(),
The 4kloc one (which I think is actually closer to 2.8kloc) doesn't include any of the SIMD ops, but it claims to implement the whole wasm spec as of 02019, and it's plausible that it actually does. A casual glance at the code doesn't reveal anything that contradicts that claim; it implements about 200 instructions.None of them include the peripherals, which are by definition excluded from wasm. But peripherals are usually the part of an emulator that requires the most effort, and they're usually a much bigger compatibility bitrot shitshow than your CPU is.
By contrast, the UM interpreter in the Cult of the Bound Variable paper was I think 55 lines of C (also, not including peripherals). My dumb Chifir interpreter was 75 lines of code; adding Yeso graphical output was another 30 lines https://gitlab.com/kragen/bubbleos/blob/master/yeso/chifir-y....
Uxn, despite its inefficiency, runs useful applications on the Nintendo DS today (in 5200 lines of pretty repetitive C, including the peripherals), and wasm doesn't and probably never will. And 365 teams in the ICFP programming contest independently implemented the UM successfully enough to run at least some existing applications; there will probably never be 300+ independent reimplementations of wasm. So in important ways they're already closer to the goal of eliminating bitrot than wasm ever will be. This is probably because eliminating bitrot isn't part of wasm's goals.
(I realize I forgot to mention the Infocom Z-machine, as well.)
Wasm has another deficiency other than complexity: as far as I know there's no standard way for a wasm program to generate some wasm and start running it, although of course the browser platform does provide that ability. This is essential if the thing you want to run under it is a virtual machine or other sort of emulator, because the only way to do an efficient emulator is to compile the code you want to interpret into the instruction set the emulator is running on. Something wasm-like can dramatically simplify this; if you're compiling to wasm, for example, you don't have to do instruction scheduling or register allocation, and you might not even have to do constant folding and function inlining.
— ⁂ —
One of the interesting ideas in the RISC-V ecosystem is that a Cray-style vector instruction set (RV64V) can give you SIMD-instruction-like performance without SIMD-instruction-like instruction set inflation. And, as the APL family shows, such vector instructions can include scalar math as a special case. I haven't been able to come up with a way to define such a vector instruction set that wouldn't be unacceptably bug-prone, though; https://dercuano.github.io/notes/vector-vm.html describes some of the things I tried that didn't work.
— ⁂ —
Why am I being so unreasonable about the amount of code? After all, a few hundred lines of C is something that you can write in an afternoon, right, so what's the big deal about 500 or 3000 lines of code for something you'll use for decades? And nobody has ever written an NES emulator in 500 or 3000 lines of code.
The problem is that, to parody Perlis's epigram, if your virtual machine definition has 500 lines of code, you probably forgot some. If a platform includes that much functionality, you have designed it so that that functionality has to live in the base platform rather than being implemented in programs that run on the platform. And that means that you will be strongly tempted to add stuff to the base platform, which is how you break programs that used to work.
In the case of MS-DOS or NES emulation this is limited by the fact that Nintendo couldn't go out and patch all the Famicoms and NESes in people's houses, so if they wanted to change things, well, too bad. NES emulator authors have very little incentive to add new functionality because the existing ROMs won't use it, and that's what they want to run.
I was trying to find the smallest Rust Wasm interpreters I could find, I should have read the source first, I only really use wasmtime, but this one looks very interesting, zero deps, zero unsafe.
16.5kloc of Rust https://github.com/rhysd/wain
The most complete wasm env for small devices is wasm3
20kloc of C https://github.com/wasm3/wasm3
I get what you are saying as to be so small that there isn't a place of bugs to hide.
> “There are two ways of constructing a software design: One way is to make it so simple that there are obviously no deficiencies, and the other way is to make it so complicated that there are no obvious deficiencies. The first method is far more difficult.” CAR Hoare
Even a 100 line program can't be guaranteed to be free of bugs. These programs need embedded tests to ensure that the layer below them is functioning as intended. They cannot and should not run open loop. Speaking of 300+ reimplementations, I am sure that RISC-V has already exceeded that. The smallest readable implementation is like 200 lines of code; https://github.com/BrunoLevy/learn-fpga/blob/master/FemtoRV/...
I don't think Wasm suffers from the base extension issue you bring up. It will get larger, but 1.0 has the right algebraic properties to be useful forever. Wasm does require an environment, for archival purposes that environment should be written in Wasm, with api for instantiating more envs passed into the first env. There are two solutions to the Wasm generating and calling Wasm problem. First would be a trampoline, where one returns Wasm from the first Wasm program which is then re-instantiated by the outer env. The other would be to pass in the api to create new Wasm envs over existing memory buffers.
See, https://copy.sh/v86/
MS-DOS, NES or C64 are useful for archival purposes because they are dead, frozen in time along with a large corpus of software. But there is a ton of complexity in implementing those systems with enough fidelity to run software.
Lua, Typed Assembly; https://en.wikipedia.org/wiki/Typed_assembly_language and Sector Lisp; https://github.com/jart/sectorlisp seem to have the right minimalism and compactness for archival purposes. Maybe it is sectorlisp+rv32+wasm.
If there are directions you would like Wasm to go, I really recommend attending the Wasm CG meetings.
https://github.com/WebAssembly/meetings
When it comes to an archival system, I'd like it to be able to run anything from an era, not just specially crafted binaries. I think Wasm meets that goal.
https://gist.github.com/dabeaz/7d8838b54dba5006c58a40fc28da9...
Performance limits what you can program on a platform, and there's a pretty wide range of interactive software where the performance bottleneck is generating pixels to put on the screen, just because there are so goddamned many of the fuckers. And fortunately that is commonly vectorizable; vector operations implemented by special CPU instructions have been a key enabling hack since BitBlt on Smalltalk in the 01970s. (In fact, the very first machine that ran a GUI had SIMD instructions in hardware, the TX-2, completed in 01958, but it didn't have pixels.)
This is even more important now than before, even without bringing GPUs into the picture. With optimized single-threaded C on an 8-core system, you're getting at best 12.5% of the machine's performance. If SIMD gives you an 8-fold speedup on slinging pixels around, which is common, you're getting at best 1.6% of theoretical performance with single-threaded SIMD-less code. With a naive, CPython-style bytecode interpreter, you have another factor of 40 or so slowdown, so you're getting 0.04% of the machine's performance: 11.3 Moore's-law performance doublings, so on today's machines you're back to 01999, age of Winamp, Quake III, ICQ, RealAudio, and Baldur's Gate. On the 02009 machine I'm typing this on you're back to 01987.
Now, people wrote a lot of interesting programs in 01987. Fractint, Hypercard, Quattro Pro, X11R3, a lot of the more elaborate Infocom games. And modern SSDs and humongous RAM mean you can do a lot of things now you couldn't do then, even if your CPU is clunking along with 01987 performance. And even a basic JIT compiler can probably get you from 0.04% up to 0.5% of the machine's performance, maybe 4% if it's multithreaded, though that makes determinism enormously more difficult.
But in theory Numpy-style vector operations can take advantage of both multiple cores and SIMD, usually giving you about 20% of the machine's performance even in a naive, CPython-style bytecode interpreter, at least on the kinds of operations where it's applicable, and without sacrificing determinism. Possibly with the kind of pipelining and tiling optimizations Numpy doesn't do you can do even better than that.
Of course, you can also design a virtual machine architecture in such a way as to permit efficient compilation, and not have to fall back on turbocharged vector operations. A lot of Wasm is designed with this objective in mind; for example, there is no dynamic typing to require runtime type checks, local variables can't alias linear memory, control flow is structured to help out the compiler, and so on, so that your interpretation/emulation slowdown is almost nil, reduced to the occasional bounds check that couldn't be hoisted out of a loop. And you can provide access to SIMD instructions, which Wasm now does, and multithreading, which it doesn't.
But all of this can't be done on a simple VM architecture. And for a given VM architecture I think an interpreter is always simpler than a compiler.
One of my stretch goals for the forever-platform is for it to be independently and compatibly reimplementable from the spec, that is, without access to a working implementation to test against, ideally within a few hours (as with Nguyen and Kay's "fun afternoon hack" desideratum for Chifir). There are indeed 300+ RISC-V implementations, but they aren't independent in the way the implementations of the UM were; the authors had not only a body of RISC-V software to work from but also other implementations of RISC-V to test the software on. But for Rosetta-Stone-level archival, we want the programmer-archaeologist to be able to get their implementation working even though they don't have an existing implementation to test against.
UM is a demonstration that this goal is achievable, if you're willing to accept that ceiling of 0.04% of native performance for the straightforward interpreter you'd write as a fun afternoon hack.
The hope with vector operations is that they give you an extra three orders of magnitude in inner-loop performance, without stretching the implementation task beyond what's achievable as a "fun afternoon hack", and without introducing so much ambiguity and so many tricky corner cases in the spec that in practice a given "ROM" will not run on an independent reimplementation. I don't know if I'll find a way to do it or not. But that's the objective.
And that's why I've been wrestling with SIMD, and don't think it's just a distraction.
So, thank you very much for the links to the relevant efforts! Oh dear, though, the V1.0 version of RVV is 111 pages, longer than the entire RV32/RV64 unprivileged spec.
> And that's why I've been wrestling with SIMD, and don't think it's just a distraction.
I think in this instance, Wasm 128b SIMD is a distraction. I agree with everything you have said, and your goals. Wasm 128b SIMD was a stop-gap and distracts from what ultimately will be a much better solution (vector processing). It is the lowest common denominator (128b width registers) impl that they could come up with to satisfy necessary immediate perf needs. It does get good speedups on code right now, but it more than doubles the number of Wasm opcodes. It isn't tenable to have the same 128 -> 256 -> 512 path that CPUs have taken. Implementation complexity and having to recode software everytime the SIMD registers change shape is lame. CPU vendors are totally ok having a needless upgrade train. The future is the past with vector processing, like RVV and flexible vectors. For archival software, I'd argue that SIMD doesn't matter. I don't think it is fare to compare RVV against the base ISA spec, the base ISA spec is trying to be the simplest possible thing, it is just a bootloader for RVV and custom instructions.
Being able to extract parallelism is way more important than constant factors, even for pathologically slow systems like Python. Thread counts for single socket systems are already at 256. One could probably implement Quake 3 in PyPy, maybe even CPython now and hit 60 FPS. CPython could definitely run Quake 1.
Wasm does support threads now. https://web.dev/webassembly-threads/
Wasm SIMD didn't even need to exist, the AST could have been encoded in a call graph that could have executed a fallback implementation, or been converted to local SIMD. It could have all just been function calls, like a binary intrinsic. It didn't really need an instruction set extension. You could encode a vector program, or something shader like in a binary wasm call graph.
I love that you use Moore Units in your comparisons, rarely do folks do this. When folks argue over 4% difference for some aspect of safety, they are talking about weeks of relative performance on a Moore Scale. But no amount of concrete performance gains will allow them to have safety or correctness. They are content to live in a ghetto all driving fast cars.
https://github.com/WasmCert/WasmCert-Isabelle
Mechanising and Verifying the WebAssembly Specification https://www.cl.cam.ac.uk/~caw77/papers/mechanising-and-verif...
However, the third cart was a metacircular interpreter, so in some sense it potentially represents a kind of reference implementation — just, one you couldn't use until your UM was running well enough to run it.
This may not be the perfect strategy but it worked well enough for 365 teams to submit a proof witness for at least one of the challenges in the Codex. And, importantly, it isn't very dependent on the hardware you're implementing the machine on, and in particular it doesn't require you to be able to run an external test suite.
— ⁂ —
You may be right about 128-bit Wasm SIMD.
When you say, "thread counts for single socket systems are already at 256", do you mean you need 256 threads to hit 100% utilization of the AVX units? Or are you talking about a chip with 64 or 128 cores with 2-way or 4-way "hyperthreading"? On https://www.tomshardware.com/reviews/cpu-hierarchy,4312.html... the highest number I see is 64 cores with 128 threads on the AMD Threadripper 3990X. Are you talking about the Tera MTA or Tilera or Altra or Parallela or something?
I'm not sure if mere vector operations scripted by a single control thread (SIMD, but in a broader sense than SIMD intrinsics, including things like RVV) can approach the level of parallelism needed.
Suppose you have 64 cores with 256-bit AVX2, like the 2.9GHz Threadripper 3990X, which Donald Kinghorn benchmarked at 1.57 teraflops on HPLinpack, and (optimistically) you're doing double-precision floating-point. In theory each core can do up to 4 flops per cycle (right?) and thus 740 gigaflops; perhaps he's counting FMA as two flops instead of one, because on the Cray-1 it would have been two flops. Without any parallelism, the max you get is 2.9 Gflops (or 5.8 if you count FMA as two), and if you have an additional 40× slowdown from naïve bytecode interpretation like CPython you get 73 megaflops, 0.01% or 0.005% of theoretical. This is maybe comparable to a Pentium MMX, Pentium Pro, or Pentium II; according to Greer and Henry's SC97 "Micro-Ops to TeraFLOPS" technical paper, matrix blocking boosted ASCI Option Red's Linpack scores from about 10 megaflops per 300 MHz Pentium Pro to about 160 megaflops per.
(This is perfectly adequate for Quake III, but maybe not at 60 fps.)
For our naïve interpreter to hit the 1.6% of maximum performance, 11.6 gigaflops, that you could get from single-threaded AVX assembly on that chip, you need an average vector length of 158 (11.6 ÷ .073). For it to approach 100% of maximum performance we need an average vector length of 10100 (740 ÷ .073). I am not convinced that you can reliably get such large vector lengths in most pixel pushing.
Even with a naïve JIT compiler with, say, 4 clock cycles per bytecode instruction, you need an average vector length of 1010 (≈1024) to get close to maximum performance, or 256 to be only two Moore generations behind (≈02018, or maybe ≈02016 given that the 3990X came out in 02020).
And the situation is 8× worse if you're alpha-compositing pixels or something; an AVX256 register holds 32 8-bit color components rather than just 4.
I think that usually you will need both SIMD and more flexible forms of parallelism to extract that much parallelism.
— ⁂ —
Why does this matter for reproducible software? Because if the forever-platform imposes a 10100× or a 1024× slowdown then we'll only use it when we really have to, and most of our software will fail to be reproducible, and we'll be back to directories full of screenshots of things we can no longer do.
— ⁂ —
Yes, 4% difference is three weeks of Moore's Law. But perhaps Moore's Law has already ended?
To, "no amount of concrete performance gains will allow them to have safety or correctness," I would add, "or durability". Modern software is not just a ghetto but a shantytown slum built from cardboard: it collapses after a few good rains.
Esp if you can operate symbolically, think triple nested loops where the inner loop is some expensive projection, if it is over a regular space, and can be decomposed as a tree. I think this language could be represented as a Wasm DAG of functions, map/pmap/apply/tree_call/reduce
Having a spec and a body of software to run but no test oracles for how that software runs feels contrived. If you have megabytes or gigabytes of code but not even screenshots of how it should it should look when it runs isn't the same as having a book in a dead language and not knowing what it should sound like.
> However, the third cart was a metacircular interpreter, so in some sense it potentially represents a kind of reference implementation — just, one you couldn't use until your UM was running well enough to run it.
This is a great example of a test oracle. Fixed-point correctness.
Say my implementation is P and the metacircular interpreter (presumed correct) is M. If P[M[X]] outputs "foo", while P[X] outputs "bar", we know there is a bug in I. This is similar to a test oracle in some ways: given a test oracle O, if O[X] outputs "foo" and P[X] outputs "bar", we know I has a bug. The difference is that in the metacircular case, the correct output might actually be "foo".
Worse, if P[M[Y]] outputs "baz" and P[Y] also outputs "baz", it does not therefore follow that "baz" is the correct output, as it would if O[Y] outputs "baz". It might just be that P has a bug that flows through to the metacircular interpreter. Maybe it rounds division the wrong way, or incorrectly considers negative zero to be unequal to positive zero.
I agree that screenshots would be very useful, but a static repository of screenshots is also not a test oracle, for a slightly different reason: you cannot submit a novel input to it and find out what the correct screenshot would be.
I agree that a DAG of functions is a useful representation that preserves significant degrees of parallelism that conventional machine-code representations expunge. However, Wasm also expunges them.
The idea of using a relational rather than functional or imperative computing paradigm for archival is interesting.
We should probably discuss this in a more amenable format.
Wasm can only run specially crafted binaries; it can't run arbitrary i386 or ARM7 code, only Wasm code.
I think the way to run anything from an era is with the following four-layer cake:
0. Whatever hardware and operating system the programmer-archaeologist happens to be running 1000 years from now.
1. An implementation of the forever-platform running on that hardware and OS. This needs to be simple enough for the programmer-archaeologist to write and debug without a running implementation to compare to, even if their hardware is base-10 or balanced ternary or has 21-bit word-addressed memory or whatever, but efficient enough to support real applications.
2. An implementation of a popular platform from the era you want to preserve, running on the forever-platform, such as wasm, MS-DOS, RV64+UEFI, or amd64+BIOS. (This doesn't have to be a platform that was widely used, just a platform to which the software you want to preserve can be compiled.) Probably you want to implement this as a binary-to-binary compiler targeting forever-platform code, not an interpreter, so that the slowdown introduced by this layer is less than an order of magnitude.
3. The application software you want to preserve, the "anything from an era".
Wasm is not even close to being a candidate for the forever-platform, but it might be a reasonable choice for (the CPU part of) the layer on top of it, because there are compilers that can compile a large body of C and C++ to it; and because it's a lot easier to compile from, and compile efficient code from, than botched monstrosities like amd64 and ARM7. (Did you realize that in ARM7 with Thumb an addition instruction can change the instruction set the processor is implementing? Because the low bit of the PC tells you whether it's in Thumb mode, and you can specify the PC as a destination register.)
In that context the fact that Wasm keeps changing isn't a big problem. You can compile some code with Wasm today and bundle it into a cart with an implementation of today's Wasm spec, and 8 or 16 years from now when you compile different code to a different version of Wasm, you bundle that code into a cart with an updated Wasm implementation.
The issue of perfection, bug-freeness, is an interesting one. You could have a specification that was bug-free but implementations that were buggy, or you could have a specification that was itself buggy. Most specification errors won't give you an unimplementable specification; they'll give you a specification that specifies behavior you didn't want.
In the forever-platform context, most such specification errors are unimportant. It doesn't matter if your bitwise operation is accidentally NOR instead of NAND, or if you accidentally specified division to round toward -∞ instead of toward 0, or whatever. As long as your implementation of the forever-platform follows the spec, and you test your forever-platform programs on the implementation and fix them when they break, they'll continue working on future implementations of the spec.
So, with respect to specification complexity, the reason the specification needs to be very short is not so that the specification has no bugs; it's so that the specification has no ambiguities that result in observably different behavior among implementations. (The spec also needs to specify something Turing-complete and reasonably efficient, but those are less difficult.)
A buggy implementation you test on would be a much bigger problem, because it could result in purportedly working "carts" (or "roms" or "programs") that in fact only operate as intended on the buggy implementation. Then programmer-archaeologists would be forced to guess what the behavior of the actual implementation was. In many cases a metacircular interpreter will not help with this: if reading outside the bounds of memory produces 0 or -1, or if division by 0 produces 0 or ∞, or if arithmetic overflow saturates or wraps mod 2³², the metacircular interpreter is likely to inherit this behavior from its host implementation.
Buggy implementations might be detectable by formal methods or by asking for alternative implementations, perhaps in a contest format as in the UM case.
— ⁂ —
Wasm3 looks very appealing in a lot of ways. 64 KiB (?) for code and 10 KiB for RAM is small enough to run on a lot of small platforms! I didn't know Wasm could scale down that small. However, most of things compiled for Wasm will need a lot more than that.
I think getting something like LuaJIT or HotSpot, or even GCC's trampolines for nested functions, working on top of the trampoline approach you suggest for Wasm, is probably technically possible but impractically slow on existing implementations.
Two things that make this impossible in a general sense: security vulnerabilities, and hardware that fails and is no longer manufactured.
It's only possible by defining the forever-platform as a sandbox abstraction, such that the containing layers can be updated for vulnerabilities and hardware.
This is what happened with the old game platforms, but even that is leaky. The NES Zapper is at least one case that remains troublesome. More generally, input and display lag always pose a problem and always will, since there always must be some translation layer for hardware that won't exist forever.
It's not just a matter of software programming. The hardware matters. For your forever-platform to truly work, it would need to include hardware manufacturing instructions from raw materials. Which is possible, but you can see where it goes out of feasible scope for any commercial endeavors.
Security vulnerabilities at that level are a non-problem. There are no security vulnerabilities in Wirth RISC, in Nock, in Chifir, in the UM, or in the λ-calculus, and there never will be. There probably aren't any security vulnerabilities in uxn/Varvara. None have ever been discovered in the 8086, which is orders of magnitude more complex than what I'm talking about. The 6502 did have some, but they were in the hardware implementation, not the architecture, and are not present in modern 6502 emulators.
Input and display lag can be a problem, it's true, but there are many applications that can tolerate a lot of lag. A guitar pedal cares a lot more about lag than a paint program, which cares more than a word processor, which cares more than a compiler.
Hardware manufacturing instructions are not needed as long as people have access to some kind of programmable computer on which to implement the forever-platform emulator.
Hardware still matters for input and output. For your platform to run a guitar pedal, you'll need some way to physically get an audio signal in and out if you want it to be of any use.
But I think it will upload to imgur, be sure to configure it off. I have it sent to have in a local folder.
It also has great .gif making support and can also screen record to a .mp4
In Windows 10, you can also press Win+PrintScreen to save a numbered screen shot to your Pictures\Screenshots folder.
On my M1 Mac I have well functioning QEMU VMs for MacOS 7, 9, 10.4, 10.11, Windows XP Windows 11 and some Linuces. I don't use proprietary SAAS formats so by all purposes I can read 99.9% of all files I have ever created with little effort.
Having screenshots of thing that are otherwise lost would just make me sad. It's like preserving a single frame from a movie and losing the rest.
I don’t think that’s necessarily true, because you are unlikely to ever rewatch the movie anyway (and then it won’t be as good as your recollection). The point is to be reminded of the fact the movie existed at a time when you wouldn’t ordinarily think of it.
Lot's of folks use us just for archiving purposes. We also have a nifty Ai that compares visual changes in the screenshots.
Your access to a website could change or the server could go down or a website may get bought.
Good reason to keep taking screenshots.
There have been dozens of times I ve been saved by a screenshot.
- insurance cards on phone - a random screenshot I took of a photo of my citizenship certificate saved me 2 weeks For travel - names of peopel - context during a call
I have a dedicated screenshots root folder with sub folders that gets synchronized with the cloud and I go through them regularly.
The problem is organizing them ...
Perhaps if I ran them through OCR, it would be easier to grep through them.
Why not convert your documents to PDF? That‘ll preserve the layout.
Mine's CloudShot. I can upload some screengrabs straightaway to Google Drive. You have the option to configure keyboard shortcuts to do that.
https://gist.github.com/soruly/889d926fcd4b2b8e8dc113a100b9d...
If you wanted to add a functionality wherein you could upload to Google Drive like CloudShot, what would you need to write?
what I wish we could record, however, is the system interactions (kinda like how you record games inside the game itself). it doesn’t record a video, but rather your mouse movements and keyboard inputs, along with the location of apps and windows and their state. it would take more space but it would be more useful in case you wanna go back and run counterfactuals.
EDIT: Ah, I realise now this achieves the same thing. But it does record video, so I'm not sure what's missing other than a visual representation of keyboard input.
1988 - Emacs 18.52
1992 - Emacs 18.59
1994 - Emacs 19.7
1996 - Emacs 19.7
1998 - Vim ????
2002 - Emacs 21.1
2006 - Emacs 21.1
2012 - Emacs 24.1
2018 - Emacs 26.1
2022 - Emacs 28.1
A spinner shows with the words Working on updates. 100% complete. Don't turn off your computer
Windows 11 is ready—and it's free! Get the latest version of Windows with a new look, new features, and enhanced security. [Download and install] [Stay on Windows 10 for now] Checking for updates ...