MicroShift
next.redhat.com
next.redhat.com
At any rate, we have a couple customers using microshift at https://kubesail.com and it works like a charm for home-hosting! Might have to add docs for it soon!
Sorry, 160meg isn't 'relatively tiny'.
I have many thousands of machines running in multiple datacenters and even getting a ~4mb binary distributed onto them without saturating the network (100mbit) and slowing everything else down, is a bit of a challenge.
Edit: It was just over 4mb using .gz and I recently changed to use xz, which decreased the size by about 1mb and I was excited about that given the scale that I operate at.
A "hello world" written in Go is 2mb to begin with so ~4mb is a bit unrealistic for any substantial piece of software written in Go. Although if that colors your opinion of Go itself, you're certainly allowed to have that opinion :)
The complex app (15k lines of code) that is distributed to these machines, written in Go, is just over 3mb compressed with xz.
Plus this is still many orders of magnitude off with trying to get 160mb off to mobile devices per the image in the article. It is a non-starter.
From the post:
> Imagine one of these small devices located in a vehicle to run AI algorithms for autonomous driving, or being able to monitor oil and gas plants in remote locations to perform predictive maintenance, or running workloads in a satellite as you would do in your own laptop. Now contrast this to centralized, highly controlled data centers where power and network conditions are usually very stable thanks to high available infrastructure — this is one of the key differences that define edge environments.
> Field-deployed devices are often Single Board Computers (SBCs) chosen based on performance-to-energy/cost ratio, usually with lower-end memory and CPU options. These devices are centrally imaged by the manufacturer or the end user’s central IT before getting shipped to remote sites such as roof-top cabinets housing 5G antennas, manufacturing plants, etc.
> At the remote site, a technician will screw the device to the wall, plug the power and the network cables and the work is done. Provisioning of these devices is “plug & go” with no console, keyboard or qualified personnel. In addition, these systems lack out-of-band management controllers so the provisioning model totally differs from those that we use with regular full-size servers.
I don't read this and think "phones". This sounds like it's targeted at embedded industrial / telecom devices. (At least based on the examples they chose, I'm sure you could use it for other things).
The word "mobile" doesn't actually appear anywhere on the page so I'm not sure where you got "mobile devices" from.
Kubernetes is a lot of code, plain and simple.
Granted fully featured app will likely use all of the module's code so it's not a factor.
Unfortunately, almost every large client library has public methods AND references someone who calls the reflection method somewhere, so in many cases you get the full tree of dependencies.
It’s unfortunate because any one dependency can bring in reflect and trigger the behavior. It hurts large generated method libraries the most since you usually only use a small subset.
See https://go.dev/src/cmd/link/internal/ld/deadcode.go line 318
staging/src/k8s.io/apimachinery/pkg/runtime/scheme.go: if m := reflect.ValueOf(obj).MethodByName("DeepCopyInto"); m.IsValid() && m.Type().NumIn() == 1 && m.Type().NumOut() == 0 && m.Type().In(0) == reflect.TypeOf(obj) {
That's quite the line of code too.
I would argue that 15K lines of code is far from a complex app as is normally understood. It's a rather small app, especially in Go which isn't exactly the most expressive language out there.
Maybe I suggest CAS (Content-addressable storage) or something similar for distributing it instead? I've had good success using torrents for distributing large binaries to a large fleet of servers (that were also in close proximity to each other physically, in clusters with some more distance between them) relatively easily.
The machines don't all need to be running the same version of the binary at the same time, so I took a simpler approach, which is that each machine checks for updates on a random schedule over a configureable amount of time. This distributes the load evenly and everything becomes eventually consistent. After about an hour everything is updated without issues.
I use CloudFlare workers to cache at the edge. On a push to master, the binary is built in Github CI and a release is made after all the tests pass. There is a simple json file where I can define release channels for a specific version for a specific CIDR (also on a release ci/cd as well, so I can validate the json with tests). I can upgrade/downgrade and test on subsets of machines.
The machines on their random schedule hit the CF worker which checks the cached json file and either returns a 304 or the binary to install depending on the parameters passed in on the query string (current version, ip address, etc.). Binary is downloaded, installed and then it quits and systemd restarts the new version.
Works great.
I think Resilio Sync might also be a good option; I think it may even use BitTorrent internally. (It was formerly known as BitTorrent Sync, not sure why they changed the name.)
I'll duplicate the comment above that the machines are all on separate vlans and don't really talk to many other machines.
The data centers don't need to talk to each other at all. We use wg at each dc for vpn access. I do need to spend a bit more time with my own wg client setup so that I can switch between dc's more easily... right now it is kind of a manual process and that is definitely a pain. One of these days...
Edit: Maybe something simple like rsync or rclone would be best for the update sync/distribution mechanism rather than BT.
That said, for now, downloading the binary really is a solved problem for me, with a relatively simple solution. I'd rather work on new features that drive revenue. =)
Same here, hence the suggestion of using P2P software (BitTorrent) for letting clients fetching the data from each other (together with initial entrypoint for the deployment, obviously), and you'll avoid the congestion issue as clients will fetch the data from whatever node is nearest, and configured properly will only fetch data outside the internal network once, after that it's all local data transfer (within the same data center).
apt-transport-ipfs: https://github.com/JaquerEspeis/apt-transport-ipfs
Gx: https://github.com/whyrusleeping/gx
IPFS: https://wiki.archlinux.org/title/InterPlanetary_File_System
Datacenter to datacenter and only 100Mb? Clearly more to the story here.... :)
You may find murder[1] / Herd[2] / Horde[3] -type tool of some use.
[1] https://github.com/lg/murder
Input binary size: 12996608
Folder contains a few more bytes for the service file and installer script.
4476005 (gz)
3421456 (xz)
3447940 (7Z)
tar c -C ./build $(BINARY) | gzip -9 - > $(PKG_NAME_GZ)
tar c -C ./build $(BINARY) | xz -z -9e - > $(PKG_NAME_XZ)
tar c -C ./build $(BINARY) | 7z a -si $(PKG_NAME_7Z)
I played around with the compression options on xz. If you have some suggestions on improving 7z, I'm all ears.Decompression time isn't an issue here.
Installing 7z on the hosts isn't great, but could be done.
> ... | 7z a -si -m0=lzma -mx=9 -mfb=64 -md=32m
3421498 (7z)
3420832 (xz)
Wild that xz is not the same every time.I just never liked the idea of taking something like open source k8s and creating a Redhat specific version that requires different treatment and whole lot of other enterprise stuff including RHEL. And it doesn't work all that better than GKE or EKS or even building your own cluster (I have done all 3.)
They should have just created tooling around standard K8s and allowed customers to use the good parts - deployment workflows, s2i etc. basically plugging the gaps on top of standard k8s. I can totally see lot of customers seeing value in that.
Huh, interesting. What do you not like about routes? My team is providing an IaaS solution for internal developers in my company and a lot developers seemes to have less problems with Openshifts service exposition abstraction (Routes) in contrast to pure Kubernetes.
It makes GitOps annoying because I don't want to treat the whole Route resource as a secret that needs to be encrypted or stored in vault.
Do I also then treat Route resources as sensitive and deny some users access on the account they could contain private keys?
I also have to worry about keeping the route updated before certificates expire instead of having it taken care of by cert-manager.
So we use Traefik's IngressRoute.
Later on, the pattern of referencing secrets in extensions became more common, and things like the NodeAuthorizer (which allows nodes to only read the secrets associated with the pods scheduled onto them) demonstrated a possible different pattern we could have chosen to implement (although nothing that can be efficiently implemented without changing kube itself today).
Agree routes should have added a ref - that was feedback informing Gateway API, and once that hits GA and provides the best of both routes and ingress, we would probably suggest using that instead. Routes is mostly frozen for all but critical new features now though so we can ensure Gateway has everything we need to replace it while still providing the necessary forward compat.
I would treat routes as sensitive. Note that within a namespace there is minimal cross user security (not part of the kube / openshift threat model), so giving namespace read access to a specific set of routes and only infrastructure users access to all routes, OR using a wildcard cert on the routers and keeping all key material out of the user’s space. On 4.x versions you could also create multiple ingress controllers and assign them to different namespaces, preventing leak between them.
This has worked well for us because not all helm charts are OpenShift friendly but they usually do allow customising the ingress resource with annotations or we patch it in via Kustomize.
https://docs.openshift.com/container-platform/4.8/networking...
We have considered having a controller that mirrors both routes and ingress to a gateway http route (since gateway is similar to routes), but plans aren’t finalized yet.
IoT is certainly a buzzword, but it also does have real meaning, and this product is aimed squarely at the IoT edge devices themselves. Seems quite appropriate to use the term IoT to describe it.
I'd love to see how the resource consumption compares to k3s.
I think one thing I'm definitely missing from the OpenShift docs [1] is reasoning. What does it add? Why do I want to learn to use an operator instead? Otherwise, it's pretty clear that it's just an operator on top of CoreDNS.
I do think that the docs are utterly devoid of Kubernetes content. I think historically, RH tried to differentiate themselves from K8s. Now, it can definitely hurt the knowledge migration and transfer.
[1] https://docs.openshift.com/container-platform/4.8/networking...
Basically all of these operators are there to allow customization. The philosophy behind OpenShift is to move everything possible to operators so they can be managed in a cloud native way. This has a bunch of benefits like being fully declarative, and able to keep your whole config in version control.
One demo using Microsoft Azure IoT hub, docker, digital twins (a glorified json from memory), and a raspberry pi was fun because it took minutes to deploy something to make a light blink.
On the naming front, why is everyone calling it MicroX. There can be only one Micro. https://micro.dev
Probably because "micro" means "small" in Greek, and the whole Kubernetes uses Greek.
> There can be only one Micro. https://micro.dev
Call me grumpy, but I spent 5 minutes on the site and still have no idea what "micro.dev" is other than "a platform for cloud-native development" (yeah, what isn't nowadays?)