56 karma · joined February 20, 2014
Socials:
https://bsky.app/profile/distracteddev.bsky.social
Interests:
AI/ML, Entrepreneurship, Fitness, Hacking, Hiking, Investment, Programming, Research, Startups, Writing, Yoga
---
1. Wait timings for jobs.
2. Run timings for jobs.
3. Timeout occurrences and stdout/stderr logs of those runs
4. Retry metrics, and if there is a retry limit, then metrics on jobs that were abandoned.
One thing that is easy to overlook is giving users the ability to define a specific “urgency” for their jobs which would allow for different alerting thresholds on things like running time or waiting.
The transition was seamless.
The Github repo mentions a REST api that can be used to push arbitrary data into the system. Are there any docs around this yet?
These people are often labelled as the "pragmatic". They also lean towards having a more general scope of knowledge and thus often lack a "speciality".
There are also those that simply wish to solve the problem in the "best" way possible, and will use whichever stack gets them closer to that ideal solution.
disclaimer: purely based on my experience/opinion/perspective.
If you do get this visa though, be very careful what you say to the border guards. You need to ensure that when speaking about your work, you always frame it as "work in support of scientific research/innovation".
The comment above (https://news.ycombinator.com/item?id=11626762) further expands on why these kinds of benchmarks, although interesting, have no real value.
Each implementation does something wildly different and responds to different inputs with completely different outputs.
To put it metaphorically, if you put a car engine in two completely different chassis and then race them on a track, you aren't gaining any real insight into relative performance of the engine in the two vehicles.
Also, just to be clear, my qualms are with the benchmarks alone, I think the library is great! Thanks for all the hard work :)
Specifically, the http server example(1), doesn't even bother using the standard library provided Cluster module(2). Cluster is specifically designed for distributing server workloads across multiple cores.
All node.js services/applications I've worked on in the past 3 years (that are concerned with scale) utilize a multi-process node architecture.
The current benchmark can only claim that a single python process that spawns multiple threads is 2x faster than a single node.js process that spawns only one thread.
This fact may be interesting to some, but is irrelevant to real world performance.
[1]: https://github.com/MagicStack/vmbench/blob/master/servers/no...
(Not to mention half of ES6 existed in CoffeeScript first, but that's a gripe for another day)
If we want to get really specific, Its also common to see the "DB" image split up between the image of the disk where the data is actually persisted, and the image of the actual DB process. This makes it easy to play around with your data under different versions of your DB.
(edit: this was a reply to the post below.. not sure how I messed that up.. gah.. such a noob)
My experience regarding building route logic into middleware is that it starts to break down as you add more complicated logic (Real-time support, webhooks, integration with a taskqueue, etc). This is mainly because the middleware pattern works best when it performs an atomic piece of common logic for a large number of routes.
However, when you start using middleware and binding it to a specific route, for a specific collection, you quickly end up with a situation where for any given route, you can no longer look at a single function and parse the flow of logic.
Also, I find these libraries deceptively simple. They may be great for your private weekend blog, but fail to address the major pain points of API implementation in a production setting:
- Configurable and flexible access control on a per-document basis.
- Rate limiting API requests
- Mutli-node deployments
- Proper, semantic, backwards-compatible API versioning
- Realtime support (Websockets, browserchannel, long-polling, w/e you prefer)
Now I'm not saying your library needs to cover all of these pain points, but I challenge you to address all of them using a middleware-based architecture while still maintaining your sanity :)Just my2c. The project looks great and hope you keep going with it and that I've provided some helpful feedback.
Cheers!
Also, whenever I've tried to use something like this for an actual product I end up quickly outgrowing the simple CRUD model since you usually need to implement additional features ontop of the CRUD logic.
By utilizing callbacks, you are free to use Async.js[0] or Step.js[1] to solve the problem you described. These libraries are great since they give you control over parallel vs series execution of the pre-requisite functions as well as solving more complex control-flow problems such as throttling, etc[2] (See link for more examples).
[0] https://github.com/caolan/async
[1] https://github.com/creationix/step
[2] https://github.com/caolan/async#control-flow
edit: Yes, you can also use similar control-flow libraries with Promises (that follow the specification) to achieve similar results but then the argument for using promises for the sake of control-flow breaks down.
https://sequelize.readthedocs.org/en/latest/docs/migrations/
Both interfaces can be abused to give you an ever growing indent and give the appearance of "callback hell"
Both interfaces can be use elegantly to help you reason about your code, make it easy to follow, and handle errors centrally.
Only one is supported natively by node.js and is the standard async interface for 90% of node.js's libraries: Callbacks.
Also, regarding "callback hell", a straw-man argument against callbacks, I highly suggest reading http://callbackhell.com/
Once you start to really scale, you'll start to find a couple parts of MongoDB break-down:
1. MongoDB has no concept of transactions and is not ACID compliant. These short-comings seem innocuous enough at first, but you end up with some crazy potential race conditions that you just end up praying you never face.
2. Distributed MongoDB layers are difficult to implement as well as maintain.
3. MongoDB Replica sets are unpredictable and unreliable in the speed at which they are able to stay up to date and require changes to your code to fully utilize (since you need to ensure you are always reading from a replica when you can, but only ever writing to the master MongoDB instance)
4. At scale, MongoDB will consistently perform worse than Postgres for CRUD-based operations
5. Do you have true relations between documents? Do you ever use Mongoose's convenient `populate()` functionality? If so, you are now making multiple db queries in series when trying to fetch a document(s) from a single collection. This starts to really hurt your query times once you get enough documents/relations in your mongo collections.
I'm sure I'm missing some, but these are some of the areas where MongoDB has fallen down for me in the past.
- The lack of an option to search only vegetarian recipes makes this pretty much useless for me.
- Adding ingredients has a fairly large delay. I suspect some optimizations could be done there.
But for some reason I'm still worried about using it as the backbone for an enterprise scale application; it almost does too much. Perhaps I'm just too averse to this amount of "magic" in any piece of infrastructure due to my Django days.
The code is open, but you're not just buying into a single module with a single purpose -- this is an entire architecture. I guess in such cases, worry is warranted before accepted it as the backbone of your company/application.
When you are paying for someone else's house, please do not be afraid to protect yourself with a proper, binding rental-agreement/contract.
Edit: I live in Canada incase that makes a difference.