Mastodon is a pretty inspirational project but the Twitter influence shows, I miss the long form writing that was encouraged before our attention spans were eroded.
Not at all close to solving it, but it’s been on my mind for a long time. Would love to hear if there are others like me out there and what you imagine such a community to look like.
Example: https://www.lfgss.com based on code you'll find on GitHub under microcosm-cc
https://github.com/ferg1e/peaches-n-stink
It's basically an experimental communication platform. Right now I am building Internet forum style communication but I want to expand to other communication mechanisms later.
You don't want the notification factory but if its a single post that gets updated you don't get any feed updates at all. Reading the same book again and again looking for the updated sections is also not much fun. Dropping a comment in the long list of comments under it doesn't really create a discussion. (specially not without feed updates)
You could design the publishing tools so that it "forces" the user into that pattern. People could work on multiple books or long reads but start with a crappy draft or just a bunch of links.
I'm planning to do a "Show HN" post next week for it and would appreciate any feedback that I could address before it. We have about 4 active users and a couple more would be great.
I wrote about the performance tuning I did in preparation to the Shown HN post: https://linklonk.com/item/277645707356438528
If you have any feedback please add a comment to that post.
> We only collect your information for the purpose of providing this service.
Okay, but what information do you collect? “Your information” is too broad; if you're collecting my retinal scans or matching me to a behavioural analytics profile, you need to justify it. If you're not (which you're not), say so!
> You can delete your account and all your data at any time (see Profile).
But can I take all my data out? Currently, no; that's not a GDPR violation, so long as you provide it on request, but it's certainly a feature I like to have!
The 30-day deletion threshold for anonymous accounts is (IANALaTINLA) GDPR-compliant, since you have to delete personal information (on request) within 30 days, and that happens even if you can't figure out whose account details you should be deleting. Good job.
I do want to add functionality to download your ratings. I'm thinking of exporting the ratings data in either the bookmarks format (ie, the format that browsers use to export bookmarks: https://support.mozilla.org/en-US/questions/1319392), csv or json. Please let me know what format would be most useful.
JSON would probably be easiest to start with, because it's easy to generate, easy to read and well-defined. Bookmarks would be a nice extra, though.
I have two questions please:
- Is there a list of all the feeds used by LinkLonk?
- How many users are there today?
To answer your questions:
1. The list of all feeds used by LinkLonk is not publicly accessible through the website. They are feeds that users explicitly submitted through https://linklonk.com/submit or feeds that LinkLonk parsed from the meta tags of the links that users submitted.
2. The number of active users has been about 4 for the last few months. I'm hoping to get it to 10 this year.
It's a shame phpbb, vbullettin and other big players in the space were too slow to adapt to mobile.
Was that a problem? My memory is that everything used 'Tapatalk' for mobile, perhaps before the first iPhone even (I recall using it on an iPod Touch).
While I understand Tapatalk have changed their business model since then, the damage was already done as the Facebook started to wholesale eat forums’ lunch as far as userbase. The biggest problem is that we never turned forum interactions into a protocol like we did with at the application level (smtp/pop, http, irc, xmpp) or on top of http like RSS, podcasts, or just plain, standardized REST APIs. This could have enabled multiple clients (like browsers) to appear and may have prevented facebooks swift dominance with online communities.
Everyone wanted to own their forum’s experience, but this stubbornness caused the fiction for users to sign up to be greater and greater. Platforms like Disqus attempted to solve this by creating an embeddable service to just drop comments in a context like a blog post, but this ultimately gave users almost no value it they were just in a shouting match against bots with generic messages laden with spam.
Facebook unified the experience for users, where the user could, with an account+app that they already had, browse and join groups and engage in discussions and become apart of communities in a way that forums could not possibly compete with.
I wasn't involved with the hosting/software/ops etc. side of it at all, but I moderated 'The Computer Forum', latterly 'Computer Juice', and used it mainly with that. I fondly remember wasting an awful lot of time helping people solve Windows problems (haven't used it since.. not saying that's related..) and spec new builds.
I suppose that's all happening on StackExchange and probably some DIY custom pc Discord server or whatever these days.
This is based in my experience:
- Old people: They started to use internet recently so they are used to social networks (Facebook, Instagram), and newspapers websites.
- Young people: Hard to make them use the browser, if there isn't an app (Instagram, TikTok), you are lost. If they want to discuss a topic, mostly Twitter through hashtags, YouTubers or Discord.
- Adult people: This is where some of them may use a forum, but you have to be lucky enough to find adult people who use internet many years ago, and they know what is a forum. If you find a 30 or 40 years old who started to use internet 5 years ago (which happens), you are lost.
And on top of that, you need to compete against Reddit and their own subreddits.
(Edited to format)
I disagree on this take.
GenZ can do long-form discussion and they use forums frequently, but for specific discussion(s).
Content consumption is done via forever-scroll apps because it's good to kill time; Tiktok, Instagram, etc match this well because it's just a stream of things to have fun with.
But GenZ is making great "longford" content on many traditional platforms, even blogging. The difference, as I understand it, is their approach to engagement. GenZ has seen Facebook arguments and said "no thank you". Twitter discourse also isn't really a thing -- people disagree and Tweet at people, but Twitter isn't like a forum thread and there's no way to ensure your content is associated with the content you want to respond to, so it's not effective to communicate in responses. Reddit is kind of a mystery for me as I just don't frequent it at all, but I don't get the impression that GenZ is posting frequently.
GenZ has platforms that work best when you make a statement, not where you open a discourse; Twitter is too fast/broad to respond to all comments and find real content to respond to, video content just isn't great for back and forth and becomes time consuming for lighter topics. Viewers will make whatever statement they want, but the validity of the video is based on how far a concept is spread; I'd actually position GenZ is very good at concisely expressing its idea in a simple and condensed format, and responses are not at an individual, but an idea. But, they will go to Tumblr/Medium/other long-form post when the medium is appropriate. This is one thing I like about a lot of GenZ content because they tend to be VERY good about choosing the right format for their argument; that many don't have much more to express besides tweets/tiktoks isn't an indictment of GenZ, it's praise. Think of the forums you maybe still lurk around and how many comments are just complete garbage/non-sequiturs; for forums such a post are enough to derail a topic/distract because we feel obligated to a degree to respond, but filtering noise is part of the skill of using more modern platforms.
Forums are kind of contrary to this, and also are bogged down by the before-mentioned Facebook-argument issue and the preferred platforms not really being strongest for direct rebuttals. Again, there are times when you'll see GenZ use forums or other long-form posts, but it tends to be more controlled or on 'forums-but-not-really' forums like Tumblr.
Forums have their purpose and use; but we have __many__ alternatives that make forums defunct, as some topics/ideas are far better expressed on Tiktok or Twitter, and whatever followup we end up with.
There's little upside to having teenagers in a forum, for example, unless you're looking to monetize.
For example: Virtual reality headsets were fairly stagnant until a young guy in his early 20s tried something new.
Old people often become stuck in their ways and it requires someone new to show up and ask how the sausage is made.
First, in my mind the difference between reddit, FB, twitter, HN, forums, etc is really just configuration. Abstracting just a tad bit higher, you can include Slack and other realtime options. I just want a curated gRPC API that implements it with pluggable auth and let others figure out discoverability and network (not activity api and not with an already built network, just persistence and auth). End to end encryption is important IMO too (even for large groups) so the host can have plausible deniability.
Second, you have to solve for hosting/network in distributed fashion without over-complicating server democratization and discoverability/naming like p2p often does. I see 2 options: 1) self hosted on at-home workstation using Tor onion service (get NAT busting for free) knowing you need an offline friendly implementation or 2) one-click easy reselling of cloud instances and domains from inside the app (this also provides a funding model).
I know of many p2p options for solving these problems, but I think we dont need to complicate at that level. As for the quality of the communities themselves, a self-hosted megaphone instead of perverse share/like broadcast incentives of companies today will automatically improve discourse (at a potential cost of creating echo chambers).
The big boys are dopamine reinforced network effects incorporated by means of software tech. So forget about aiming for Homer Simpson.
You'll have to be happy with a small, but productive minority. Enough valuable people would rather die than use FB for niche purpose X. Start by convincing them...
[1] https://gemini.circumlunar.space/ Earlier discussion at: https://news.ycombinator.com/item?id=23042424
Sooooo..... StackOverflow solved it by starting with a vertical the founders had a lot of social juice in, and spreading in to other verticals. Possibly also by focusing very tightly on "questions and answers".
So my suggestion is "overfocus". If the big platforms have a weakness, it's that they're generic one-size-fits-all solutions. Solve one problem (Q&A, show off my project, discussion, news aggregator) for one vertical really well, then expand. An example off the top of my head might be collaborative note-taking for a class. 6510's "write a book in public" platform is also a fantastic idea at first glance.
(But skip the distributed bit - customers don't care about your architecture, and whatever USP a distributed platform has can be emulated by a centralised platform. Centralised always wins).
> Centralised always wins
My dream isn’t necessarily to win in a financial/monopolistic sense, but rather to build a compelling enough alternative to the centralised systems that have lost their way thanks to incentives that aren’t aligned with the community.
Facebook, reddit, disqus all started out with good intentions to connect people, but have been slowly eroded by incentives to suck user attention.
So it may not the best business strategy, but I think such software should live or die on whether the community enjoys using it and is willing to (financially) support the continued existence, rather than how much attention can be siphoned into ads.
In other words, small niche communities where a few members don’t mind contributing financially rather than huge communities that rely on network effects and centralisation.
Ok, so. The problem you're trying to solve is "build a community that is prepared to contribute financially to the running of the site" (correct me if I'm wrong).
Distributed is one possible solution to this problem. You're in love with that solution and its time to murder your darlings. Sit down, brainstorm five other possible solutions, and honestly assess which one solves your problem best.
I think you'll have a hard time beating a subscription model.
One of the main problems with platforms like FB and Reddit is that the posts/discussions are shortlived. They bubble up to the top in the feed when they're fresh and active but then die off and are replaced by the next thing that craves your attention.
Forum posts are sorted in chronological order, grouped into categories. Browsing through the feed you can see what topics are being actively discussed, or you can search for past discussions on a topic you like, and resurrect and old thread if you find a good one, and it gets a new life. I like this model.
One concept I've thought about is something like Reddit, where anyone can create and manage subs, but which wouldn't have the same kind of karma whoring and short attention span issues, i.e. posts/threads would live forever and not get locked after 6 months or whatever Reddit does, and posts would be sorted by activity rather than ADD points. I've found so many good X years old Reddit threads with super interesting discussions which I would've liked to jump in and resurrect but can't.
Of course an immediate problem that comes to mind is spam. If posts with new comments are lifted to the top it invites spam bumps and thus moderation. And/or it could be combined with some sort of karma system (reputation, account age etc).
And, since this wouldn't be a specialized forum you'd need to make sure you could cater to various different kinds of communities, i.e. have good multimedia sharing capabilities (for those communities only wanting to share images/videos/memes), code formatting and syntax highlighting, maybe LaTex etc...
Basically, every leafy green (and herbs, and even mushrooms), can grow in a range of climatic condition (phenotype, roughly) ie temperature, humidity, water, CO2 level, pH, light (spectrum, duration and intensity) etc. As you might have seen around the world there is a rise in indoor vertical farms, but the truth is that 50% of those are not even profitable. My startup wants to discover the optimal parameters for each plant grown in our indoor vertical farm and eventually I would let our AI system control everything (something like alphaGo, but for growing plant X (lettuce, kale, chard, ). Think of it as reinforcement learning with live plants! I am betting on the fact that our startup will discover the 'plant recipes' and figure out the optimal parameters for the produce that we would grow. Then, the goal is that cities can grow food cheaper in more secure and sustainable way than our 'outsourced' approach in country side or far away lands.
So now I have secured some funding to be able to start working on optimizations, but I realized that *hardware* startups are such a different kind of beast (I am a good software product dev though, I think). Honestly, if anyone with experience in hardware related startups (or experience in the kind of venture I am in) would just want to meet me and advise me, I would take it any day. Being the star of the show, it's hard for me to handle market segmentation, tech dev, team, next round of funding, European tech landscape, etc. I am foreseeing so many ways that our decisions can kill my startup, all I need is advise from someone qualified/experienced enough. My email: david[at]hexafarms.com
At the very least what's a link to your startup's website?
If you want users to have it from your profile, put it in the “about” field.
Sounds similar to what I read a long time ago about a big tomato farm in the Netherlands... Have you tried talking to actual farmers of that produce? Universities? Agricultural faculties do a lot of research in that direction.
Expensive, quickly perishable produce might be able to compete, otherwise I guess free water and energy from above in the "remote" classical farming will be hard to beat.
And then my naive guess would be that to generate enough data for a "ml" approach not only by name might be somewhat expensive.
This sounds so negative, but this is not my intention... I wish you all the best and hopefully will stumble upon a success story in the future :-)
I heard about a greenhouse company that has programmed their climate control to match “best growing conditions historical weather”. So, they ask local experts what year / location had the best X and then they use that region’s historical weather and replay it in their greenhouse. I thought that was brilliant!
(Just realized this was Kimbal Musk that mentioned this)
This had replaced shortening hormones in modern gardening (or at least at that greenhouse, but my understanding they were just doing the same thing as everyone else).
I guess there is a lot more to learn for those who have scale enough to experiment and patience to follow through.
I agree that micronutrient content has decreased in the past century. Some might be because of scale, some might be that yield gains are mostly driven by macronutrients and water, not micronutrients, it could be selecting varieties that taste better, or it could be depleting the soil.
That said, the US has an obesity epidemic, so there's no shortage of macronutrients. Macronutrient shortages also seem rare. Scurvy and rickets aren't exactly problems.
I currently have a UI with the comments down the side of the screen which looks like this:
https://www.volt.school/videos/c980297a-417b-416f-947b-58a70...
This is good because you can easily: - See all the comments - Navigate between them - See replies etc.
However it has a huge problem with you trying to balance watching the video with reading the comments.
I also have an alternative UI I've been working on which only shows one comment at a time:
https://www.volt.school/videos-v2/c980297a-417b-416f-947b-58...
However the downsides of this is that you can't see all the comments at once. I'm not a UI/UX designer AT ALL so I'd really appreciate some pointers around how to think about making this better! The original post mentions "close to solving", I think I am pretty close but it's still not quite right and while I'm not out of ideas yet, I'd appreciate feedback if solving this is obvious to someone else.
How about breaking the play into chapters/zones/rooms/segments (whatever makes sense for the game) then showing all the comments for that segment. Once the segment ends, there would be a replay button if they missed anything on the first play while reading comments, and a next segment button to carry on.
Interesting time spans could be marked for slow motion, boring bits played at double time.
There would be high level navigation between segments with thumbnails and comment counts. Buttons to skip between “pivotal moments”, maybe with voting to highlight them.
Something like Soundcloud comments but for video.
Asian video platforms used to do that. Here's an example: https://www.youtube.com/watch?v=hOMMQmYwd4I
It's totally crazy but can be made much more coherent. Would be useful to have comments at a specific time and place on the screen just for very accurate pointers/comments.
I don't know how normal my use is though or if that's at all helpful.
EDIT2: to further comment on current table architecture, we have 3-4 other tables with minimal fields (3-4 Boolean/Char fields) that are relationally linked back the Videos table with a char field ‘video_id’, that is unique on the Videos table. Again, not a proper foreign key so no indexing.
if youre having issues with 5gb, you will face exponential problems when it grows due to lack of indexing
If you post those queries here, we can probably give tips on how to improve the situation.
Tangentially related for those who have experience, I am using Django-silk for latency profiling.
If you write a small program to check the integrity of your blobs i.e. that the structure of the json didn't change over time, you may be able to infer a relational table schema that isolates those bits that really need to be blobs. Too leave it too long invites long term compatibility issues if somebody changes the structure of your json objects.
If you have any foreign key columns, add indexes on them. And if you’re doing any joins, make sure the criteria have indexes.
Similarly, if you’re filtering on any of the nested JSON fields, index them directly.
This alone may be sufficient for your perf problems.
If it isn’t, then here’s some tips for the blobs.
The JSON blobs are likely already being stored in TOAST storage, so moving them to a new table might help (e.g. if you’re blindly selecting all the columns on the table) but won’t do much if you actually need to return the JSON with every query.
If you don’t need to index into the JSON, I’d consider storing them in a blob store (like S3). There are trade offs here, such as your API layer will need to read from multiple data sources, but you’ll get some nice scaling benefits here and your DB will just need to store a reference to the blob.
If your JSON blobs have a schema that you control, deprecate the blobs and break them out into explicit tables with explicit types and a proper normalized schema. Once you’ve got a properly normalized schema, you can opt-in to denormalization as needed (leveraging triggers to invalidate and update them, if needed), but I’m betting you won’t need to do any denorm’ing if you have the correct indexes here.
And since you have an API layer, ideally you’ve also already considered a caching layer in front of your DB calls, if you don’t have one yet.
First of all, I think the caching layer (which we currently don’t have) is going to be a necessity in the coming weeks as we scale for an additional project (that will be relying on this architecture)
Second of all, it is just PK lookups. We don’t actually have a single fk (contractor did not set up any relations), which makes me think moving all of this replicated JSON data from fields to tables may help.
The queries that are currently causing issues are not filtering out any data but returning entire records. In ORM terms, it is Video.objects.all(), and from a URL param in our GET to the api, limiting the amount of entries returned. What’s interesting is this latency scales linearly, and at the point we ask for ~50 records we hit the maximum raw memory alloc for PG (1GB) causing the entire app to crash.
The solution you propose for s3 blob store is enormously fascinating. The one thing I’d mention is these JSON fields on the Video table have a defined schema that is replicated for each Video record (this is video/sensor metadata, including stuff like gps coords, temperature, and a lot more).
So retrieving a Video record will retrieve those JSON fields, but not just the values: the entire nested BLOB. And does so for each and every record if we are fetching >1
Would defining this schema with something like Marshmallow/JSON-Schema be a good idea when you mention JSON schemas we control? As well as explicitly migrating those JSON fields to their own tables, replaced with an FK on the Video table?
Regarding JSON schema, if you have a Marshmallow schema or similar, yes that’s a wonderful starting point. This should map pretty closely to your DB schema (but may not be 1-to-1, as not every field in your DB will be needed in your API).
I’d suggest avoiding storing JSON at all in the DB unless you’re storing JSON that you don’t control.
For example, if the JSON you’re storing today has a nested object of GPS coords, temperature, etc.. make that an explicit table (or tables) as needed. The benefits are many: indexing the data becomes easier, the data is stored more efficiently, the table will take up less storage, the columns are validated for you, you can choose to return a subset of the data, etc… You will not regret it.
In either scenario, I’d still generally encourage avoiding storing JSON(B) unless there isn’t a better alternative. There are a lot of maintenance, size, I/O, and validation disadvantages to using JSON in the DB.
Or at least as you suggest if required for performance the data would still be stored denormalized and where needed materialized / document-ized?
At my current company, there seems to be a belief that everything should be moved to mongo / cosmo (as document store) for performance reasons and moved away from sql sever. But really I think the issue is the code is using an in house orm that requires code generation for schema changes and probably less than ideal performance query generation.
But then I am also aware of the ease of horizontal scaling with the more nosql orientated products, and trying to be aware of my bias as someone who did not write the original code base.
As a general rule of thumb, yes. Starting with denormalization often opens you up to all sorts of data consistency issues and data anomalies.
I like how the first sentence of the Wikipedia page on denormalization frames it (https://en.wikipedia.org/wiki/Denormalization):
> Denormalization is a strategy used on a previously-normalized database to increase performance.
The nice thing about starting with a normalized schema and then materializing denormalized views from it is that you always have a reliable source of truth to fall back on (and you'll appreciate that, on a long enough timeline).
You also tend to get better data validation, reference consistency, type checking, and data compactness with a lot less effort. That is, it comes built into the DB rather than introducing some additional framework or serialization library into your application layer.
I guess it's worth noting that denormalized data and document-oriented data aren't strictly the same, but they tend to be used in similar contexts with similar patterns and trade-offs (you could, however, have normalized data stored as documents).
Typically I suggest you start by caching your API responses. Possibly breaking up one API response into multiple cache entries, along what would be document boundaries. Denormalized documents are, in a certain lens, basically cache entries with an infinite TTL... so it's good to just start by thinking of it as a cache. And if you give them a TTL, then at least when you get inconsistencies, or need to make a massive migration, you just have to wait a little bit and the data corrects itself for "free".
Also, there are really great horizontally scalable caching solutions out there and they have very simple interfaces.
I’ve already done the legwork, cloning to the current prod DB locally and playing around with migrations, but the fear of applying anything potentially-production breaking is scary to a dev who has never had to work on a “critical” production system!
1. What are the CRUD patterns to the "blobby data". 2. What are the read patterns and how much data needs to be read.
Until Read/Write are properly understood the following solutions should be considered as general guide lines only.
If staying in PG: JSON can be indexed in Postgres You could also support a hybrid JSON/Relational model giving the best of worlds.
Read:
Create views into the JSON schema that model your READ access patterns and expose them as IMMUTABLE Relational entities. (Clearly they should be as light weight as possible)
Modify:
You can split the JSON blobs into their own skinny tables. This should keep your current semantics and facilitate faster targeted updating.
Big blobby resources such as video/audio should be managed as resources and not junk up your DB
Warning:
Abstracting the model into multiple tables may cause its own issues depending on how you ORM map your entities.
Outside the Box Thinking:
-Extract and transform the data for optimized reading. -Move to MongoDB or a Key Value store
Conclusion:
What are the update patterns? -is only 1 field being updated -inter-dependencies of the data being updated How are "update anomalies" minimized
You will need to create a migration strategy to a more optimal solution and would do well to start abstracting with views. As the data model is improved this will be a continuous process and the data model can be "optimized" without disturbing the core infrastructure requiring rewrites.
As someone said, indexes are the best to do lookups. Remember your DB engine do lookups internally even if you are not aware (joins, for example), so add indexes to join fields.
Another thing that worked me for me (and I dont know if it's your case), was to add trigram text indexes, which make it faster to do a full text search. Remember, anyway, that adding a index makes search faster, but insert slower, so be careful if you are inserting a lot of data.
Rearchitecting the schema might be worth doing. From the technical side, PG is pretty nice about doing transactional schema changes. I'd be more worried about the data though. Are you sure that every single row's Json columns have the keys and value types that you expect? Usually in this type of database, some records will be weird in unexpected ways. You'll need to find and account for them all before you can migrate over to a stricter schema. And do any of them have extra unexpected data?
- Change the field type from JSON to JSONB (better storage and the rest) https://www.postgresql.org/docs/13/datatype-json.html
-Learn the json in-build functions and see if one of them can replace one you made ad-hoc
- Seriously, replace json with normal tables for the most common stuff. That alone will speed up things massively. Maybe keep the old json around in case, but remove when it become old(?)
- Use views. Views allow to abstract over your databaee and allow to change internals
- If a big things is searching and that searching is nkind of complex/flexible, add FTS with proper indexing to your json then use it as first filter layer:
SELECT .. FROM table WHERE id IN (SELECT .. FROM search_table WHERE FTS_query) AND ...other filters
This speeeeeedupppp beautifully! (I get sub-second queries!)- If your query do heavy calculations and your query planner show it, consider move them into a trigger and write the solved result into a table, and query that table instead. I need to loans calculations that requiere sub-second answers and this is how I solve it.
And for your query planner investigation, this handy tool is great:
An hour with the query planner can save you days or weeks or wasted work!
Related: I'd love to have an Android app with a shortcut that allows me to quickly translate Google Maps links into coordinates, OSM links or other map links. There is a browser extension that does this on desktop, so if anyone is looking for a low hanging fruit idea for an Android app, this might be a fun idea (if I don't get around to it first).
Everything is seamless for me, though admittedly I'm not a super heavy calendar user.
I plan to do a write up on my whole Google-free setup, but I haven't done it yet, unfortunately.
It's called Code Shelter:
It's stalled for a while, so I don't know how viable it is, but I'd appreciate any help.
For example, I use a javascript library that's best in class for what it does, and yet hasn't had any real commits from its maintainer since 2016. There are 50 pull requests open, some of which fix significant bugs, or add good new features. There are literally 2000 forks of the library, some of which are published on npm but are themselves unmaintained, and almost none of which link back to the actual fork's code from npm. It's a mess, and I bet it's a situation repeated hundreds of times over.
If you were to figure out a workflow by which a maintenance team could form on your platform, and then a) the existing maintainer is pinged to request that they add the team, falling back to b) making it easy for the new team to fork and adopt existing pull requests while supporting them through initial team-forming by laying out a workflow for assigning needed-roles, then I think you'd have a valuable platform.
The key thing is ensuring there's a large enough team to start, so that yet another fork doesn't die on the vine, so maybe think about a(n old) reddit link type interface where people can link to, vote on, and volunteer for projects, with no work needed until there's critical mass and the platform moves the project forward.
I certainly understand the rationale but doesn't this narrow down the universe of possible maintainers while putting even more load on existing maintainers by expecting them to take on more work?
Hadn't heard that malapropism before.
I haven't found any implementations I'd consider good. The problem as I see it is that there are tree based algorithms like https://webspace.science.uu.nl/~swier004/publications/2019-i... and array based algorithms like, well, text diffing, but nothing good that does both. The tree based approach struggles because there's no obvious way to choose the "good" tree encoding of an array.
I've currently settled on flattening into an array, containing things like String(...) or ArrayStart, and using an array based diffing algorithm on those, but it seems like one could do better.
This is ever more important with the onset of remote hiring, remote work, and the isolation/depersonalization it brings to newcomers to the industry.
There's also an "evil" momentum in remote hiring -- some companies _need_ asynchronous interviews to support their scaling and operations, and the general perception is that it's impersonal and dehumanizing.
This made me think that if we preemptively answered interview questions publicly, then it'd empower the job seekers to have a better profile/fight back a dehumanizing step, while allowing non-job-seekers to share the lessons that were important to them.
I've been getting decent feedback on my attempt at the solution HumblePage https://humble.page, the reality is that there's a mental hurdle to putting your honest thoughts out there.
One feedback about the homepage: show a few examples of how people have answered questions, below the prompt. That's more helpful to get us thinking about our own answers, compared to a blank field. (Also, it's not clear what the percentages are meant to represent there. And I'm guessing the number next to the edit icon shows how many people have answered the question already? May need some UI tweaks on these.)
My intention was to show the general comfort level of answering the prompt in public. Looking back, I wonder if I was being too quirky.
I’m thinking the same on needing UI tweaks, I’m planning for major rearrangements.
Thank you for your interest. Please feel free to reach out via the contacts if you’d like an invite.
The closest most known example of this kind of game nowadays is Roblox, but I'm thinking of things more like Mario Maker or the older-generation Atmosphir/GameGlobe-likes.
Unlike "modding platforms" or simulators/sandboxes/platforms such as Flight Simulator, VRChat or ARMA, these games' content are created by regular players with no technical skill, which means the game needs to provide both the raw assets from which to build the content, as well as the tool to build that content.
Previous titles tried the premium model (Mario Maker), user-created microtransactions (Roblox) and plain old freemium (Atmosphir and GameGlobe).
I suspect Mario Maker only works because of the immense weight and presence of the franchise.
Roblox's user-created microtransactions (in addition to first-party ones) seem to be working, but they generate strange incentives for creators, which I personally feel taints a lot of the games within it. (The user-generated content basically tends to become the equivalent of shovelware)
GameGlobe failed miserably by applying the microtransaction model to creator assets, which means that to make cool content, creators have to pay as well as spend lots of their time actually building the thing, which means most levels actually published end up being the same default bunch of free assets and limited mechanics.
Atmosphir is a bit closer to me so I find some more nuance in its demise, but long story short, essentially they restricted microtransactions only to player customization, however it didn't seem to be enough to cover the cost of developing the whole game/toolset. Eventually adding microtransactions to unlock some player mechanics, which meant that some levels were not functional without owning a specific item.
---
In short, the only thing that can effectively monetize on is the game itself (premium model) or purely-cosmetic content for players. Therefore, to incentivize the cosmetics, the game needs to be primarily multiplayer, which implies lots more investment on the creator tooling UX, as well as the infrastructure itself. But this also restricts the possibilities for the base game somewhat.
Core is very similar to Roblox in that the creation tools are rather involved, it tends more to a platform with distinct creator/consumer roles.
There's also the PS4 game Dreams, as well as other integrated-modding initiatives like in Krunker.
The solution is a general-purpose distributed computing platform designed for end-user development.
The closest three things that exist are Google Sheets, replit.com and dndbeyond.com. Replit is too low level, dndbeyond is not powerful enough, sheets are stuck with grid and too clunky for everything else.
Here's a few things the user should be able to do:
1. Design a tabletop roleplaying game system from scratch and automate all the math
2. Write content designed to be used with a system
3. Use systems and content designed by other people, without copy-pasting
4. Modify the system and the content designed by other people for your own purposes
5. Share access to the content in a granular way
Tabletop roleplaying games are unique: they thrive on content that must be created and shared quickly, but includes simple but fully general programming capabilities. Seems like a great place to start making programming as commonplace as literacy.
Planetside 2 has very slight pay to win mechanics in the form of subscriptions for more xp, but it doesn't feel bad to play without pay.
Counter-strike on the other hand (and I think so other valve games too) have just about the perfect model in my mind. There are no advantages you get by paying, only status. The skins that you can buy and sell look cool but are purely cosmetic.
Even with this in mind people spend quite a lot of money (we're talking hundreds of dollars for one gun skin in some cases). It always seemed like a great way to generate revenue ethically.
One thing I will note with that model is they still have the gambling mechanics with the "crates" that open random skins. You could probably crank it up one more ethical notch by getting rid of those or trying to make them less addictive.
However, paid status is not really a big hook for single-player games, you need to be able to show it off! This means that the game must be designed primarily around multiplayer interaction, which is fine but limits a lot the kinds of games you can implement with this monetization.
Does anyone know of any books or surveys about statistics and medicine, or specifically mechanics of the human body?
- One therapist is taking measurements of say your arm motion and making inferences about the motion of other muscles. He does is very intuitively but wants to encode the knowledge in software.
- The other one has an oral appliance that has a few parameters that need to be adjusted for different geometries of the mouth and airway.
The problems aren't that well posed which is why I'm looking for pointers to related materials, rather than specific solutions (although any ideas are welcome). I appreciate replies here, and my e-mail is in my profile. I asked a former colleague with a Ph.D. and biostats and he didn't really know. Although I guess biostats is often specifically related to genetics? Or epidemiology?
I guess the other thing this makes me think of is software for Invisalign or Lasik, etc. In fact I remember a former co-worker 10 years ago had actually worked on Lasik control software. If anyone has pointers to knowledge in these areas I'm interested.
This seems like a sequential Bayesian filtering problem. Probably high enough dimension that you should just use a particle filter. The big seminal background text in this area is Bishop: Pattern Recognition and Machine Learning.
If the "motion of other muscles" is inferring pose, you could also look into what computer graphics calls inverse kinematics (a typical IK model has a number of dimensions that could fit into a particle filter). There's some more in-depth stuff in motion planning that actually takes into account muscle capability. But I wouldn't know where to find info on that, short of watching the last several years of Siggraph Technical Papers Trailers, grabbing all the motion planning ones, then reading everything they cite.
I've heard of inverse kinematics but I think it's more focused on "modeling" than statistics/probability? That is, you would have to model each muscle?
I think he is doing something that is more "invariant" across human variation? (strength, body dimensions, age, etc.) I'm not sure which is why my question was vague, but this is helpful.
Hard to know whether that's relevant without knowing what he's trying to predict though.
For the oral appliance adjustment, I'm not sure what your output measures of interest are. If they're mechanical maybe you want to do a sensitivity analysis using FEA. Maybe look at FEBio: https://febio.org/
As for books or surveys, biomechanics is huge topic so I'm not sure what to recommend without wasting your time. If you're still defining the problem, maybe run some searches on Pubmed with the "review" and "free full text" boxes checked, and browse the results until you find which sub-sub-topic is relevant to you?
https://pubmed.ncbi.nlm.nih.gov/?term=biomechanics&filter=si...
If no one on the team knows statics, dynamics, and (if you're considering internal strain and stress) continuum mechanics, consider finding a mechanical engineer to help.
I think the basic idea is that when you're doing physical therapy that targets certain muscles, you have to find the muscle(s) that are limiting the motion! This is not obvious because they all interact.
Like if you have a back problem, you can try to exercise your back all you want, and that may not actually fix the problem. Because the real issue could be with your leg, which causes 16 hours a day of "bad" motion against your pelvis, which in turn messes up your back.
All the muscles in the body are interlinked and they often compensate for each other. When people have a problem in one area, they compensate in other ways.
So I have the same question as above: I think inverse kinematics is more about "modeling"? You would need to model every muscle, which is hard, and it is specific to a person?
I think his intuition is partly based on a mental model, but it's also probablistic. I think the model has to capture the things that are "invariant" across humans (i.e. basic knowledge of anatomy), and the variation between humans is the probablistic part. It's also based on variation in your personal health history / observed behavior, e.g. how you walk, how often you're sitting at a computer, etc.
So it does feel like an "inference" problem in that sense -- many factors/observations that result in multiple weighted guesses of the cause / effective therapies.
Inverse dynamics takes the kinematics data from above and, in combination with ground reaction forces measured from a force plate (or instrumented footwear, etc.), calculates the forces and moments on each joint. Since control of the human musculoskeletal system is over-determined (the same motion, forces, and moments can be produced by multiple muscle activation patterns), EEG data or even ultrasound elastography is sometimes used to better constrain estimates of muscle activation patterns.
In your example the usual approach would be to use (elements of) the above methods to find out if a patient had unusual motion patterns, like the suspected abnormal leg motion in your back pain patient. Statistics comes into play once you have population data to classify as "good" or "bad", and when you're trying to determine if the hypothesized relationships between symptoms and particular motion / muscle activation patterns genuinely exist. Of course, it's fine to try different approaches (but don't forget to obtain IRB review and comply with the various regulations on human subjects research).
Imagine that a product description is a n-dimensional vector like:
( manufacturerName, modelName, width, height, length, color, ...)
Now imagine you have a file with m such vectors (where m is in millions), and that not all fields in the vectors are reliable info (typos, missing info, plain wrong, etc).What is a good way to determine which product descriptions refer to the same product.
Is this even a good approach? What is state of the art? Are there simpler ways?
Here is what I mean by good:
- robust to typos, missing info, wrong info
- efficient since both m and n are large
- updateable (e.g. if classification was done, and 10k new descriptins are added, how to efficiently update and avoid full recomputation)(edit: I am at a desktop now and I can say a bit more)
Here is the process in a nutshell:
1. Create a fast hashing algorithm to find rows that might be dups. It needs to be fast because you have lots of rows. This is where SimHash, MinHash, etc. come into play. I've had good luck using simhash(name) and persisting it. Unfortunately you need to measure the hamming distance between simhashes to calculate a similarity score. This can be slow depending on your approach.
2. Create a slower scoring algorithm that measures the similarity between two rows. Think about a weighted average of diffs, where you pick the weights based on your intuition about the fields. In your case you have handy discrete fields, so this won't be too hard. The hardest field is name. Start with something simple and improve it over time. Blank fields can be scored as 0.5, meaning "unknown". Hashing photos can help here too.
3. Use (1) to find things that might be dups, then score them with (2). Dump your potential dups to a CSV for human review. As another poster indicated, I've found human review to be essential. It's easy for a human to see that "Super Mario 2" and "Super Mario 3" are very different.
4. Parse your CSV to resolve the dups as you see fit.
Have fun!
hamming_dist = bin(a^b).count("1")
It relies on a string operations, but takes ~1 microsecond on an old i5 7200u to compare 32bit numbers. In python 3.10 we'll get int.bit_count() to get the same result without having to do these kind of things (and a ~6x speedup on the operation, but I suspect the XOR and integer handling of python might already be a large part of the running time for this calculation).If you need to go faster, you can basically pull hamming distance with just two assembly instructions: XOR and POPCNT. I haven't gone so low level for a long time, but you should be able to get into the nanosecond speed range using those.
I'd start by detecting common typos. Typos are similar to un-typo'd data, so I'd do a frequency analysis on the textual representations of manufacturer name and model name, and a Levenshtein distance calculation, then synonymise the obvious synonyms (looking things up when I wasn't sure). The key idea is that you have access to more information than just this dataset: Tony and Tomy are different manufacturers, but Sony and Somy aren't (even though somy is in the dictionary and tomy isn't).
Once the manufacturer and model fields are mostly typo-free (after typo replacement – don't modify the original file, if you can help it!), you can start looking at dimensions and colour. Sort by manufacturer, and start de-duping entries. Once you get a feel for the process you're doing (e.g. under what circumstances do you check whether there's a 102mm Phillips screwthread?), you can start automating bits of it. There will always be special-cases, but your job is to get the data processed correctly, not to get the computer to process the data.
Accidentally aliasing two different products is much worse than leaving the same product described twice, so err on the side of “these are different”. (Keep in mind that manufacturers of some things, e.g. SD cards, often pretend two different products are the same – so you can't always win!) Remember, humans exist: bothering them a few million times is a problem, but bothering them a few hundred would be okay.
When new data comes in, I'd run all the code I used to come up with my system, and see if the output was notably different. If it was, I'd get the computer to let me know.
I'd also add some way for users to flag duplicates. Many humans make light work.
I built a commercial system like that for Thermo Fisher, except their descriptions were encoded as natural language text on input, not vectors (for an extra complication).
Some observations:
1. Crude methods based on vector embeddings, cosine similarity, Levenshtein, etc – don't work, if you care at all about false positives.
I see sibling comments recommend this, but it's clear this cannot work if you think about it. Values like "black" and "white", or "I" and "II" (part numbers), "with" and "without", are typically close together in such crude representations, but may lead to products that are not interchangeable.
2. A hybrid approach worked. The SW produced suggestions for which products might be duplicates (along with a soft confidence score), then let a human domain expert accept / reject these suggestions. It also learned from these expert decisions as it went, to save human time.
What I quickly learned is that even as a human (programmer with a PhD in ML), I could not look at two product descriptions and make the decision myself. Are these the same product or not? One word, even one letter, could be absolutely vital. Or absolutely irrelevant. Sometimes even the same attribute / word, depending on the product category.
Hence the final interactive solution with a domain expert in the middle. It worked well and saved time, rather clever, but not in the "hooray NN training" way. A lot of work went into normalizing the surface features intelligently based on context: units, hyphens / tokenization, typos…, because that's a mess in product sheets. The "fancy" downstream ML and clustering part was relatively simple by comparison.
But YMMV, the Thermo Fisher products were fairly specialized and sophisticated (in their millions).
Many sourcers and recruiters don't have a technical background and find it very difficult to hire software engineers, especially in the current labor market which is very tight.
I'm starting off simple: writing recruiting guides from a software engineer's perspective that are easy to understand.
Are there other ways we can make technical recruiters better?
- list salary range for positions
- emphasize tech stack
- emphasize number of rounds
I’ve wasted my time in the past going through the interview process only to find out the company’s budget for the position was only up to X after getting the offer; a valuable lesson I’ve learned to avoid since then of course.
I also see some recruiters only talking about what the business does … leaving out the tech stack.
If these points are clear and easily visible to recruiting leads they might get higher quality candidates.
Just my two cents. ¯\_(ツ)_/¯
Basically, take the whole “how I learned my data science skills” into something that can be done in public.
The recruiters can then see a wide range of examples, and can be better at picking up where people’s strengths and talents are.
(This is focused on analytics)
Is there a way to use multiple threads or GPU to encode pngs? I haven't been able to find anything. The images are 3500x3500px and compress from roughly 50mb to 15mb with maximum compression (so don't say to use lower compression).
Another neat feature is that it's designed to be progressive, so you could host a single 10mb original file, and the client can download just the first 1mb (up to the quality they are comfortable with).
Take a look: https://jpegxl.info/
Although these don't directly solve the PNG encoding performance problem, maybe some of these ideas could help?
* if users will be using the app in an environment with plenty of bandwidth and you don't mind paying for server bandwidth, could you serve up PNGs with less compression? Max compression takes 15s and saves 35MB's. If the users have 50mbit internet, then it only takes 5.6s to transmit the extra 35MB, so you could come out 10s ahead by not compressing. (yes, I see your comment about "don't say to use lower compression", but no reason to be killed by compression CPU cost if the bandwidth is available).
* initially show the user a lossy image (could be a downsized png) that can be quickly generated. You could then upgrade to a full quality once you finish encoding the PNG, or if server bandwidth/CPU usage is an issue then you could only upgrade if the user clicks a "high-quality" button or something. If server CPU usage is an issue, the low then high quality approach could let you turn down the compression setting and save some CPU at the cost of bandwidth and user latency.
Now zlib specifically is focused on correctness and stability, to the point of ignoring some fairly obvious opportunities to improve performance. This has led to frustration, and this frustration has led performance-focused zlib forks. The guys at AWS published a performance-focused survey [1] of the zlib fork landscape fairly recently. If your stack uses zlib, you may be able to find a way to swap in a different (faster) fork. If your stack does not use zlib, you may at least be able to find a few ideas for next steps.
[1] https://aws.amazon.com/blogs/opensource/improving-zlib-cloud...
Not terribly hard if you only need 1-2 formats supported, e.g. RGBA8 only. You don't need to port the complete codec, only some initial portion of the pipeline and stream data back from GPUs, the last steps with lossless compression of the stream ain't a good fit for GPUs.
If you want the code to run on a web server, after you'll debug the encoder your next problem is where to deploy. NVidia teslas are frickin expensive. If you wanna run on public clouds, I'd consider their VMs with AMD GPUs.
If you don’t care about cost of ownership, use CUDA. It only runs on nVidia GPUs, but the API is nice. I like it better than vendor-agnostic equivalents like DirectCompute, OpenCL, or Vulkan Compute.
Maybe you could write the png without compression, compress chunks of the image in parallel using 7z, then reconstitute and decompress on the client side.
[1] https://blender.stackexchange.com/questions/148231/what-imag...
Here's what that custom format might look like.
(I'm guessing these images are gray scale, so the "raw" format is uint16 or uint32)
First, take the raw data and delta encode it. This is similar to PNG's concept of "filters" -- little processors that massage the data a bit to make it more compressible. Then, since most of the compression algorithms operate on unsigned ints, you'll need to apply zigzag encoding (this is superior to allowing integer underflow, as benchmarks will show).
Then, take a look at some of the dedicated integer compression algorithms. Examples: FastPFor (or TurboPFor), BP32, snappy, simple8b, and good ol' run length encoding. These are blazing fast compared to gzip.
In my use case, I didn't care how slow compression was, so I wrote an adaptive compressor that would try all compression profiles and select the smallest one.
Of course, benchmark everything.
We are on JIRA now, and it’s … JIRA. We tried basically any other tool, including Excel (yes, that is somewhat possible).
My problem generally is that tools are slow, planning is cumbersome, visibility is limited and reporting for clients is often even more limited.
Heck, I’d even write my own tool if I knew it would help others, but I am concerned it’s too close to what we already have for anyone to actually migrate.
You could help me by sharing your thoughts!
You could also modern try agile tools, for example Linear. JIRA is good for 100+ teams and complex architectures.
Not affiliated, but I've had a positive experience with it in a small team. I would describe it as an IDE for issues.
https://clickup.com https://clickup.com/on-demand-demo
ClickUp for Agile Workflows https://www.youtube.com/watch?v=H9hZRwivnL8
I'm also working on a conversational search engine (using NLP) at http://supersmart.ai
I am starting to consider alternative tools such us wireguard to reduce load, but I am concerned of adding too much complexity. Tinc's mesh network makes setup and maintenance easy. The wireguard ecosystem seems to be growing very quickly, and it's possible to find tools that aim to simplify its deployment, but it's hard to see which of these tools are here to stay, and which will be replaced in a few months.
What is the best practice, in 2021, to ensure all communication between cloud VMs (even in a private network) is encrypted?
DIY: envoyproxy.io / HashiCorp Consul for app-space private networking over public interfaces.
LowCode: Mesh P2P VPN network among your clusters with FOSS/SaaS like WireTrustee / tailscale.io / Slack Nebula.
Outside of the WireGuard ecosystem there's ZeroTier [2] which has been around for a while and they're working on a new version; and Nebula [3] from Slack, which is likely to be maintained as long as Slack uses it.
There might be others, but with tinc these four are the ones I've seen referred to most often.
Have you noticed whether it is worse for lots of small requests vs large data transfers?
I use a very similar setup, but haven't seen tinc CPU usage matter yet, though for very low traffic.
Given that you can study Introduction to Computer Science from Harvard University, online, for free and in your own time, it seems like the barriers to building skills is lower than ever.
However, many people are put off or intimidated by the idea of studying such a course. My solution to this is some kind of mentoring, either 1-to-1 or more likely in small groups. However, this is very resource intensive for my idea to scale. I'd be very interested to hear how others might approach this, both the mentoring or the underlying encouragement to study.
Say you're writing a novel. Every writing session (but the first, obviously) should cover the end of the last scene and the beginning of a new one, AND THEN STOP, i.e. not finish the new scene.
Your brain will want to come back to the work to finish it, which overcomes the friction of "starting" something new every time.
It's easier said than done. It's surprisingly difficult to leave something unfinished at the end of each work session. But that's the trick.
Starting is easy.
Thinking about new things and maybe throwing down a few ideas is indeed easy and pleasant.
But deciding to spend a few hours to move a project (instead of not) is what the brain hates. It hates commitment, and is very afraid of the opportunity cost.
It's quiet so early in the morning, so your productivity will skyrocket. I've coded my Paras this way while working a full time, heavy blue-collar job.
Here's an example: there's a lot of ML tutorials on doing image identification. Like you have a series of images: picture one might have an apple and a pear in it, picture 2 might have an apple, orange, and a banana in it.
Where I'm struggling is putting this into my domain. I have a 100k images and from that around 1k distinct labels (individual images can have between 1 to 7 of these different labels), with between 13,000 to 100 images as examples in each label.
Is that enough data? Should I just start working on gathering more? Is this a bad fit for a ML solution?
With regards to the size of your dataset, there's no hard rules it depends on the complexity of the task. Problems with a high number of classes are among the more difficult ones, 100 samples per class might be enough or it might not . The only way to know for sure it to try and see if you reach a performance that's acceptable for your application.
I recommend the Pytorch framework, it's coherent, easy to use and well documented (both the API reference and the examples available on SO and throughout the net). Your problem is similar to imagenet (assuming you want to detect the presence of a class and not its position, in which case it's a different problem.) so you can try to run one of the pytorch tutorials and see how well it does. The only difference is you want to detect multiple classes in one single image so you'll have to adjust your output layer and loss potentially but the network itself could remain the same. You also might want to look into doing transfer learning with a pretrained imagenet network to speed-up the training.
1k distinct labels with a long tail distributions which is what you are describing, is definitely a challenging problem. It's called a imbalanced classification problem.
I'd focus first test how well your model will be able to predict these classes doing a stratified cross validation (stratification controlled by the label class), and measure the F-score, Weighted Accuracy and ROC AUC. Check also the precision and recall for each class. You'll definitely see that the model predicts better for the labels with more samples. The code that you use here, you'll be able to reuse later on, so keep it organized and easy to follow.
Then you have a couple options, focus on gathering more examples for the labels with small sample size, or try to oversample your dataset. This article is a good place to start https://towardsdatascience.com/4-ways-to-improve-class-imbal...
Considering this problem of image classification is normally solved with deep learning, the more data you have, the better will be your results.
Second, this could be more than enough! Especially if you are doing transfer learning.
Third, you can "inflate" the amount of images you have now with "Image Data Augmentation"
It is a good fit for ML, but you need to be clear on what the results will be. If you expect 100% accuracy, that won't happen. Even 90%+ accuracy would be require a lot of effort.
This is expected to enable us to solve distributed coordination problems. Also, it should facilitate richer more meaningful relationships between people.
Expected outcomes include increased thriving and economic productivity.
[edit: consider the limit on how many people you can know and the relationship between how deeply you come into relationship with that population and the size of that number]
[edit: i.e. count positive and negative sentiment statements assigned to speakers and compare the per speaker ratio to the experimentally determined minimum "healthy" ratios not yet replicated]
You're right that there needs to be a tractable starting place. This is not lost on me. I may have used a flexible definition of "close to solving" but one's interpretation also fits into the scope of the effort. I'm at least 10% into it! ;P
An hypothetical solution would be a system that talked a language similar to plain english, but that was determinist. You let people write their problems and views to the system, and the system determines which are the widest consensus available within a given scope and what are the highest priority problems (perceived by people). This has a lot of problems, but it's a good way to think about the topic. Even with such a system, would you really be solving the problems you want to solve?
If it does, then this is basically symbolic AI. You can try to relax requirements... but you kinda need an "automatic coordinator". If you go with a manual coordinator instead, then I doubt you will be able to scale anything that's not extremely rigid and hierarchical, at which point you are re-introducing many of the same problems you were trying to fight in the first place.
I think the direct human input method is given too much focus although it and related interactions have their place. The fallible sensors directly reporting readings from reality already has sufficient noise related issues. I suspect more richly informing people will yield better results.
I am inspired by stories such as the fish farm pollution problem [0]. Consider how a reality based game theoretic analysis of agent choices might guide your selection of future work mates (or lakes) and facilitate a different friction in finding your next contribution to the world.
[0] search "3. the fish" on https://www.lesswrong.com/posts/TxcRbCYHaeL59aY7E/meditation...
>> A combination of "all categories are fuzzy" and "all models are wrong but some are useful"?
Are you talking about my first paragraph or symbolic AI?
>> The fallible sensors directly reporting readings from reality already has sufficient noise related issues.
I assume here you are trying to say that human input is not reliable.
I don't understand what's your approach with AI here. You seem to want to use it to better inform people? How? You are going to say that human input is not reliable, but then train an AI that can't explain itself and expect people to take its advice? Either noise can be palliated at scale in both places or none.
Finally, I'm very familiar with meditations on moloch. But you seem to be betting on an "education-based" solution, which doesn't fit very well with the scenario that meditations on moloch exposes, which is not that some people couldn't make better choices (for society, the collective), but rather that the "questionable" choices of a few can deeply compromise the game for everyone else. I mean, we all probably agree that it would be great to educate people on these concepts, but I doubt that will be enough to stop the dynamics that cause it.
> Are you talking about my first paragraph or symbolic AI?
In the link you provided and the second paragraph of your first reply you seem, to my reading, to suggest using a system to facilitate discovering agreement on specific actions, knowledge, and tactical choices. Stated differently agreement within groups, perhaps large groups. You discussed in both comments the challenge of being specific and static, which is the downfall, in my opinion, to many symbolic systems - the presumption that our ability to discretely describe reality is sufficient. To me fuzzy categories and useful broken models comment about that finding. The systems you are describing sound useful but seem to solve a different problem than I mean to target.
> I assume here you are trying to say that human input is not reliable.
Yes, I find human output to be unreliable and I believe it is well understood to be so. An example of a system that has elements of scaling social knowing is Facebook. I believe it is well understood that people often (and statically speaking prevalently) present a facsimile of themselves there when they are presenting anything actually more than superficially adjacent to themselves at all. This introduces varying amounts of noise in to the signal and displaces participation in life, perhaps in exchange for reduced communication overhead. Humans additionally make errors on the regular, whether through "fat fingers", an unexamined self, "bias", or whatever. See also "Nosedive" [0].
> I don't understand what's your approach with AI here
I haven't really described it - the ask was literally for the problem, not for solutions. There is a certain level of vaporware in my latest notion for exactly how to solve it. As stated obliquely however, there are aspects of the solution that I don't really want to be dragged through a discussion on here on HN.
> an AI that can't explain itself
I haven't specified unexplainable AI. I actually see evidence based explainability as a key feature of my current best formulation of a concrete solution. That, in context presents quite a few nuts to crack.
> Finally, I'm very familiar with meditations on moloch
I only meant to link the fish story but the link in MoM was broken and I failed to find a backup on archive.org, not putting a whole ton of effort into looking.
Consider how the described "games" change if those willing to cooperate and achieve the maximal outcomes could preselect to only play with those who are inclined to act similarly? If you grouped the defectors and cooperators to play within their chosen strategies based on prior action? Iterated games have different solutions and I find those indicative of life, except that social accountability doesn't scale. In real life such specificity is impossible and no guarantees exist. Yet, I believe that the rights systemic support structures could solve a number of problems, including a small movement of the needle towards greater game theoretic affinity and thereby a shift in the local maxima to which we have access.
Most reference to Wikipedia are dead links.
Many legacy media will stealth edit articles or outright delete them.
Original media files can be loss and after strange eons their authenticity will not be able to be asserted.
It will soon be impossible to distinguish from deep fakes and actual original and genuine media.
Some regimes such as Maoist China wanted to rewrite their past from scratch and erased all historical artifacts from their territory.
There are strong pressure to create an Orwellian Doublespeak to erase certain words entirely from speech, books and records. With e-books now the norm it has now become legitimate question to ask if the books are the same they were when the author published them.
Collaborative content websites have shown that they were not immune to subversive large and organized influence operations.
I have set my mind to multiple solutions (even bought a catchy sounding *.org domain name!). Obviously it will have to be distributed as to build a consensus and thus it will have to rely on hashes. But hashes alone are meaningless so some from of information will have to come along with them, which in themselves are information to authenticate with other hashes. I was thinking that the authentication value would come from individual recognized signatories. Those would be a a mesh of statements of records. For example you might not trust your government, but you might trust you grandparents and you old neighbors who all agree that there was a statue on the corner of the street and they all link to each other and maybe link to hashes of pictures and 3D scans with links. Future generations can then confirm those links with other functional URIs.
Something like blockchain technology seems an obvious choice but I have no experience with that (for now) but also there is the problem that it needs to be easily usable; therefore there is a need of a bit of a centralization (catchy domain name yay!) although any one could setup his/her own service for certain specialized subjects.
Thoughts?
Given these snapshots, you could write manual extractors or build a machine learning system [1] to extract the main content of each page as plaintext. Then load these timestamped text file snapshots into git, which will give you a hash of the content and let you easily track changes.
Push the git repo to a few places like github, bitbucket, and maybe IPFS where people can mirror it.
[1] https://joyboseroy.medium.com/an-overview-of-web-page-conten...
Alternatively you could use the built-in reader mode in Firefox or Chrome which do this automatically, but then you'd have to figure out how to maintain stability of the extraction algorithm between new browser releases.
A huge bonus there would be when the order difference can be represented in a graph, so that tesselation or other approaches like a hypercube representation can be used for quick estimations. (that's what I'm aiming for right now)
If successful, the next step would be to integrate it into my web browser so that I can try out whether the equilibrium works as expected on some niche topics or forums.
My problem is that it doesn't matter how I design the thing, either the screws offer too little precision so I can't help but to crush the tip into the sample every time, or too little travel distance so I can't help but crush the tip into the sample when adjusting the coarser screws near the tip. This is the kind of thing that looks like a non-problem on the web, because everybody just ignores this step.
Perhaps you could somehow attach the piezoelectric component or bars to a micrometer [1] which is designed for accurate and repeatable measurement?
Yet, they are a bit expensive. I'm still not willing to budget all that, but I'm starting to consider it.
[1] https://www.aliexpress.com/wholesale?catId=0&initiative_id=S...
Looks like the exact thing I need are micrometer heads, and some even come with nice threaded mounts.
https://www.global-optosigma.com/en_jp/Catalogs/pno/?from=pa...
If you have machining/fabrication skills, it might also be possible to buy a few worm gear sets and modify your micrometer to move really slowly but precisely.
I've got to try this on my next attempt.
In the future, if you'd like even more precise measurements, theoretically you could use 2 different frequencies or a reflected source, and look at the interference or superposition of the waves.
I'm by no means an expert in this, but I've heard that optical measurement (eg: laser + Michelson interferometer) could theoretically take you down to the nanometer range.
But it's easy to go overboard with this, haha.
https://www.osapublishing.org/oe/fulltext.cfm?uri=oe-20-5-56...
https://iopscience.iop.org/article/10.1088/0957-0233/9/7/004
Here's an open source stage project using flexures that will likely help
https://openflexure.org/projects/blockstage/
Also, see Dan Gelbart's 18 video series about building prototypes
But it's a great assembly for anything that doesn't have a feedback on the positioning.
Any help would be greatly appreciated.
Hawking radiation ?
Laser tunnel ?
Magnetic canon ?
Centrifugal launcher ?
Vacuum diffusion ?
Electrical beam lensing ?
That actually does make for a fairly efficient (better than fusion in energy per unit fuel mass) reactor design in principle, but you need a sub-solar-mass black hole (in the 1 billion - 100 billion ton range), and there's no known practical way to produce one.
... Not being flippant; I find that these kind of prompts can help in thinking from a new approach.
Of course google can't do it. But this is a ripe for someone to step in.
We are going to start work on this in a weeks, so I'm looking for some insights/shortcuts/existing projects that will make our lives easier.
The goals is to process events from students during exams (max 2500studnets/exam = ~100k-150k events) and generate notifications for teachers. No fancy ML/AI, just logic. Latency of max 1 min.
Our current plan is to let a worker pool lock onto exams (PG lock) and pull new event every few seconds for those exams where (time > last pull & time < now - 10s). All the notifications that are generated are committed together with a serialized state of the statemachine and the ID of the last processed event. Events would just be stored in PG.
This solution is mean to be simple, be implemented in really short timeframe and be a case study for a more "proper & large scale" architecture later on.
Any tips, tricks or past experiences are much appreciated. Also, if you think our current plan sucks, please let me know.
Having had a few cracks at this problem, in my opinion using locks is the wrong approach.
What you will want is:
* split all input data in batches (eg batches of 10k records, or periodic heartbeats every X seconds, etc)
* assign each batch a unique identifier
* when writing data to the output store, store the batch id along the data;
* when retransmitting a batch for whatever reason, reuse the same batch id and overwrite any data in the output store that matches this batch id.
Obviously this becomes more tricky when you’re dealing with eg window functions or more complex aggregations.
In this situation, I believe that an approach such as “asynchronous barrier snapshotting” works best. Every X seconds, you increment an epoch. While incrementing, you stop ingestion. Then you first tell the output source to create a checkpoint, then the input source to create a checkpoint, and once both have been checkpointed, you can continue streaming data again.
Anyway, these are two approaches I’ve used over the years that work well. Explicit locks don’t work well in distributed processing, imho.
What is the peak and avg QPS you need to support? High peak QPS might force you to introduce distributed workers and makes locking impractical.
Another consideration is how much do you care about data integrity. Would it be a problem if a few messages are lost? What if a message for processed twice? What if servers lost connection to db for a few seconds? What if a whole server/db goes down?
It’s a more future proof than building this on top of Postgres.
You can use txid_current_snapshot() and friends to track the last "timestamp". Proper use of locks will help you avoid the complexity associated with long-lived transactions.
Exactly-once semantics can be tricky to guarantee if you do it at the wrong layer of abstraction. Sometimes building exactly-once semantics on top of at-least-once semantics is the way to go.
Kafka and rabbit MQ are both overkill under 100 events/sec. The extra ops overhead isn't worth it. Besides, with PG it'll be nice to be able to always query a couple tables to completely discern the state of the system.
I have a corpus of text in many Indian languages, which i'd like to index and search. The twist is that I'd like to support searches in English. The problem is that there are many phonetic transliterations of the same word (e.g the Hindi word for law can either be written as "qanoon" or "kanun"), and traditional spelling correction methods don't work because of excessive edit distance.
My approach is this: Use some sequence to sequence ML technique (LSTM, GRU, ..., attention) to a query in English to the most probable translation and then use that to look it up using a standard document indexing toolkit like Lucene. (I can put together a training dataset of english transliterations of sentences to their original text)
The problem is that I'd like the corpus, the index and the model to be all on a mobile. I have a suspicion that the above method won't straightforwardly fit on a mobile (for a few Gig of corpus text), and that the inference time may be long. Is this assumption wrong?
How would you solve the problem? Would TinyML be a better approach for the inferencing part?
Is the corpus the only data you have, i.e. do you need to use it for training and validation as well?
In terms of the size of the data, if you want to store the corpus on the phone anyway, won't the index and model be relatively small in comparison?
It is not necessary that there is a one to one correspondence between words. Sometimes two english words may represent one hindi word, or vice versa.
I believe I can build up a decent sized training/validation set, for example from Bollywood song lyric databases written in English, and mapping them to the Hindi equivalent (or Tamil, Bengali, etc).
As for your last question, I dont know, since I haver implemented an ML model in practice. I saw a tutorial on Bert this morning, where a word has 768 features. That itself sounds huge, leave alone the model itself.
There is an algorithm called Soundex with python implementations you can try.
The fundamental difference between my problem and traditional spelling correction algos is that in the latter, there is a canonical correct spelling to be used as a reference. In my problem, there isn't. There are different approximate ways of spelling out most hindi words ... there is no one correct way. There are common patterns, sure, but it is too tedious to encode all the variations.
I designed a custom STT model via Kaldi [0] and hosted it using a modified version of this server [1]. I deployed it to a 4GB EC2 instance with configurable docker layers (one for core utils, one for speech utils, one for the model) so we could spin up as many servers as we needed for each language.
I would recommend the WebRTC or Gstreamer approach, but I wouldn't recommend trying to build your own model. It's really hard. Google's Cloud API [2] works well across lots of accents and the price is honestly about the same as running your own server. If you want to host your own STT (for privacy or whatever), I'd recommend using Coqui [3] (from the guys that ran Mozilla's OpenSpeech program). Note that this will likely be much, much worse on accents than Google's model.
[1]: https://github.com/alumae/kaldi-gstreamer-server
[2]: https://cloud.google.com/speech-to-text
Edit: Forgot to mention, there's also a YC company called Deepgram that provides ASR/STT as a service, you could give them a shot: https://deepgram.com/
I discuss some intellectual problems and solutions.
The process might span different medium (write email, do something in the app, check twitter, etc) and different activities multiple days. How to make sure they know what they should do next? Checklist? Emails? Slack? Wizard?
My problem is how to bill people for consuming object storage properly. Do you do it retrospectively and take the fraud risk? Are there any pre-existing platforms that do Ceph billing?
https://www.digitalocean.com/blog/why-we-chose-ceph-to-build...
I’m working on community discussion boards which exist at the intersection of interests.
Eg. Mountain Biking/New Zealand, Propulsion/Submarine, Machine Learning/Aquaponics/Tomato, etc.
The search terms for interests are supplied from Wikipedia articles which avoids duplicate interests and allows for any level of granularity.
I find that key word functionality in search engines has degraded to the point that finding good content for niche interests is difficult. I’m hoping with this system I can view historical (and current) discussions around my many niche interest combos.
I’ve got the foundation done, I just need some feedback/advice on whether I’m reinventing the wheel here, or if others share this problem?
And the goal of this isn’t to replicate those more sweeping discussion boards because they’re great!
My issue is that once things get more niche, the subreddits are tough to find - they could have obscure names, and even just learning about them often involves another user recommending. Plus creating a subreddit for every niche intersection isn’t ideal.
With this I could zone in on exactly what I’m after, without sifting through all the unrelated stuff. With the Aquaponics/AI/Tomatoes, I’d be dealing with only that intersection point. Not the peripheral stuff.
I’m considering moving the binary data into S3 and then doing the sync layer on the server (which means the front end requests the data from the backend and is given it back as a JSON object with base64 values).
Doing this manually via code isn’t impossible, just API intensive, so I’m wondering if this is a solved issue for anyone.
The why: The JSON blobs are recordings of words and sentences that can be copied between articles.
Where/why is your current system failing/inadequate/cumbersome?
Why do you want to move the data to s3?
Why are delete cascades important?
For various reasons -- mainly familiarity, FOSS ecosystem, and cross-platform compatibility -- I'm going to try to implement this in Free Pascal / Lazarus. There is one kind of component I'm definitely going to need, and if there were a ready-made one I could use in stead of building one from scratch, it would save me a lot of time and effort. I've looked around online, but so far haven't found "the perfect one". So, my question is:
Can anyone recommend a good FOSS graphical SQL query builder component -- i.e, one which presents tables and their columns to the end user, so they can specify joins and filters by clicking and dragging, etc -- for Lazarus (or in a pinch Delphi, to port to FP/L)?
Reasoning: helping with the code behind a paper on explanatory AI systems.
Related code on github if curious https://github.com/pollomarzo/map-generation/tree/main/graph
EDIT: thanks for suggestions will give them a spin throughout tomorrow :)
Iframes with postmessage() where needed (like dynamic window size changes) isn't pretty, but it's easy to do.
Users should be able to install whatever software they want. Similar for developers, they should be free to publish whatever software they made.
Apple/Google approach is suboptimal because centralized point of failure. And they do censorship on their stores, both political and arbitrary.
Linux approach is suboptimal because users don’t have keyboards to create these sources.list text files. Even if they had qwerty keyboards, I don’t like the UX, too hard to use.
Traditional P2P like bit torrent + DHT is suboptimal because smartphones, would use too much electricity + bandwidth to be practical.
So far, I’m thinking about developers-hosted binary packages, and existing code-signing infrastructure for authenticity and integrity (Verisign, Comodo, Digicert, these guys, up to developers to choose one). The configuration issue from Linux should be solvable with QR codes scanned by camera, plus a custom URI handler for the web browser on the phone.
The main thing I don’t like about that approach — a store app on the device is a good UX from end users’ perspective. Yet it seems impossible to make one with that approach.
I’m very far from being blocked on that yet, but I will face that problem eventually.
P.S. I’m not going to solve security at that level. As Android store shows us, it’s borderline impossible even with Google’s resources. Modern mobile SOCs have enough juice to solve that properly, on the lower levels of the stack. Most of them support hardware-assisted virtualization. All of them are fast enough to run proper multi-user Linux, with security permissions and SELinux kernel module.
1. Ritualize
I notice a pattern with my kids: going to bed is a ritual, and any deviation is reason enough for them to leave the bed.
2. Slow down in advance
Going to bed after playtime is impossible. So cut screens, playtime, play/listen some quiet music, read books or newspaper (again, no tablet/reader), at least 1h before bed time.
3. Recap the day
Remind your kids of their day, activities and make them aware of the fatigue. Works better when in bed, with mine.
4. Stay with them if they're afraid
Learn why they're afraid, teach them why there's no reason to be afraid. I've had to hang a sock to the door every night for months to scare tigers away :D It just works ^^
Every parent knows the pain and every kid has their own back story and the relationship with the parent(s) is key to finding a way into bed.
Eventually, they will sleep.
In our case, we settled that unless exceptional situations, our kids had to fall asleep in their own bed, because we wanted/needed our intimacy. To get that, I had to stay in my kids' room for as long as 2hrs for months, but didn't let go. Today, going to bed is thankfully not a situation anymore.
I'm not sure this is in any way helpful, but here's my shared experience and learnings. YMMV.
So for us, when I first started to do this. Each night they get a 'treat' but to get that treat they first need to be ready for bed - eg bedroom ready to sleep in, correctly dressed/ washed etc..
then after the treat they must choose a calming activity - ideally in their bedroom eg reading (nothing that gets their heart rate up) for 30-60mins then they must bush their teeth
It's this point we say time for bed, but we allow them to carry on reading for another 30-60min then it's lights off
if they don't do the activities/ actions after the treat, then we warn them that they'll not get one tomorrow etc. (and really do what you say)
also you may need to flexible on the activities until they get into the swing of it
(The only thing worse than trying to get a child who isn't tired to go to sleep is a child who is too tired to go to sleep.)
Apparently I also still have a sense of humor. Or maybe I don't, because perhaps pretending one doesn't get the joke of doubling down on xkcd silliness is perhaps a joke in itself, which I didn't get.
Also, bedtime trouble usually means not enough outside time and physical activity during the day. Or that the kids want more of the parent's time.
I realize that some parents reach for it every night and this is not something I'm suggesting.
https://www.amazon.com/Bringing-Up-B%C3%A9b%C3%A9-Discovers-...
1. make sure they don't fall asleep with something they cannot keep all night (i.e. while you are singing to them, rocking them, sitting next to them or when they are drinking a bottle of milk etc)
2. make sure they understand that even if you leave the room it is just temporarily. Small kids are - for good reasons - very afraid of being forgotten or left alone.
2.1 Using a timer to remember to visit the room regularly and often as they learn to sleep alone can help a lot
2.2. Increase the interval each day. I increased it by two minutes each day.
2.3 If the kids are happy in their bed, continue to visit their room at the scheduled time: you don't want them to think that you forget them if they don't cry.
Using this method I've got my last few kids to enjoy going to bed and sleep better in less than a week fo each of them.
If there's someone trying to use CV to recognize stuff and we are trying to prevent it (basically, a black box situation) - is it viable to use adversarial attacks at all?
Will it work long term? Can they overcome our AA by downsampling and adding noise? Can we make another AA that would still prevent them using CV? If it's better to chat elsewhere - my Twitter is in my profile.
Thanks, it didn't update my profile for some reason, fixed!
This includes entirely different ways of visualizing free trade, etc.
I have made some significant progress in the past year, but I am running into this headache where the papers I need are not easily accessible (or cheap) since I'm not an academic and I'm literally doing this as a for fun side project. I've been trying to find a university or college willing to take me on in some capacity simply to let me get access to academic catalogues, but I am not having that much luck unfortunately.
The best results recently though include some really fascinating stuff that comes from statistical measures that give clarity to how a region would respond to being isolated for any period of time.
Same question for effects systems, where would I look to understand how they're designed and the trade offs for their design decisions?
Erlang's messaging may have potential as another alternative
Since the Linux kernel has pluggable schedulers, that code might be well-structured for reading. Again, I don't know the specifics.
I would recommend not to start with coroutines in the beginning if your main focus is scheduling. In the end coroutines and async/await is about building a userspace scheduler on top of a scheduler that already exists in the OS, so you just get twice the amount of logic. However the schedulers used in userspace are often a lot more trivial than the OS ones, since they don't support preemption or priorities. Erlang might be the exception and an interesting thing to look into.
The system is a bunch of batch jobs that are scheduled to run at different intervals. These jobs can be modelled as an acyclic directed graph of steps. They basically download files from vendors and map the rows inside them into a generic format (for generating reports). There are a lot of vendors and each vendor can have a different file format containing different fields -- hence requiring custom business logic to populate (map) the corresponding generic file (like aggregating fields, fetching values from DB, etc.). Also these vendors' files sometimes contain errors, or are dropped late for download, etc. -- failures can happen and these failed instances of jobs should be able to rerun.
Existing system is built using Spring Batch and Spring Integration. The problems with the existing system are:
1. there are more than 200 jobs and most of them have their own custom logic during mapping -- cannot be generified
2. lot of manual work needed to onboard new vendors
3. jobs are synchronous and run only on one node, typically for lots of hours
4. rerunning jobs is a nightmare
Dream state for this system:
1. Dynamically add jobs to the runtime using generic components that can be reused -- maybe through an API / UI
2. Preferably, multiple records from a single file be processed across distributed nodes to generate a single output generic file
3. Rerunning should be easier
I am a noob to CS. I did a good bit of research for the past month. Found a few data-science tools in Python -- which is a no-no for a production system. Also, I know that the steps cannot be made generic after some extent since custom mapping logic is required for almost every vendor. But asking to see what is possible. Any help to point to prospective tools and technologies to solve the above will be much appreciated.
Thanks
The problem: how can I make decisions based on a sample with a binary question. I think the central limit theorem applies, and I need to account for various priors and missing votes. Is there an existing solution to this problem? The server is written in nodejs, if that matters.
This applies when our questions are not just trying to learn about the world (e.g. 'Our survey discovered that 10% of posts are considered misleading'); we are going to use their answers to decide on actions, e.g. removing posts, attaching warning labels, etc.
Those answering the questions know this, and (if they have a preference over which action will be taken) are incentivised to give more extreme answers. A classic example is an ice cream company surveying shoppers about the flavours they like: if I truthfully answer that I like chocolate slightly more than strawberry, this will have a small effect on the survey result, and hence the company's new product flavours. However, if I falsely say that chocolate is the best flavour I've ever encountered, and that strawberry makes me vomit, that will have a much stronger effect on the survey result, and make it more likely that the company will make the chocolate ice cream that I prefer.
The "bayesian truth serum" counteracts this by asking each question in two parts: there's the initial question we want answered, as well as an additional question: "how do you think others will answer?". For example:
- "I find this misleading" and "I think 80% of respondents will find this misleading"
- "I rate strawberry as 4/5" and "I think 10% of respondents will give strawberry 1/5; 20% 2/5; 50% 3/5; 15% 4/5; and 5% 5/5"
The first answers (the ones we care about) are weighted based on two conditions: how closely the estimated distribution matched the real answers, and how 'surprisingly popular' the first answer is.
To see why this cancels-out the incentives to lie: our best chance of affecting the result is to choose a 'surprisingly popular' answer, since this will contribute more weight to the result. However, these two constraints exactly cancel out:
- The answers we predict are popular, will also be those we predict are unsurprising (after all, we could predict them!)
- The answers we predict will be surprising, will also be those we predict are unpopular (that's why it would be surprising if they were popular!)
It turns out that the rational strategy, for swaying decisions as much as possible towards the outcomes we want, is to answer the first part truthfully.
A similar analysis applies to answering the second question (the estimates) truthfully. In that case there are two things to consider:
- We want our estimates to be as close as possible to the true distribution, in order to maximise our response's weight.
- We want to engineer our estimates such that the answers we disagree with get a high estimate, and hence appear 'unsurprising' (reducing the weight of those responses). Our estimates must sum to 100%, so decreasing the 'surprisingness' of one answer must increase the 'surprisingness' of the others. The effect we have on each answer's weight will be small, but it will affect every response which chooses that answer. Hence to have the largest impact, we need to decrease the 'surprisingness' of those answers we think will get the most responses. Yet that exactly what we've been asked for (an estimate of how popular we think each answer will be!)
> (if they have a preference over which action will be taken) are incentivised to give more extreme answers.
With yes and no answers, how can answers become more extreme? If you are asked "Is this misleading? Yes/No" and it's only marked as misleading when "Yes" is the majority, then you are incentivized to answer with your true opinion. If you want the post to be marked misleading, then answering yes increases the chance that it is marked as such.
The bayesian truth sereum information score makes sense when you are trying to reward people for truthfully answering your questions; for example by paying them [0]. When asking "Is this misleading?", how do you use the information score to compute who won?
[0] - http://www.eecs.harvard.edu/cs286r/courses/fall10/papers/DW0...
However, if I know my answer will affect censorship, etc. then I may try to predict the resulting distribution, and vote yes if I predict it's less than 75%, and no if I predict it's more than 75%.
For example, I may be more "trigger happy" if I think people are more likely to believe something uncritically; I may be more of a "devil's advocate" if I think something is under-represented, or less likely to be taken seriously.
> A similar analysis applies to answering the second question (the estimates) truthfully.
How does this avoid (or compensate for) downweighting the preferences of people who are legitimately ignorant about what everyone else thinks (and consequently give estimated distributions that hardly match the real answers at all)?
Although truth-telling is a Nash equilibrium of this setup, it's not the only one. However, as α → 0 the truth-telling equilibrium becomes dominant (i.e. achieves a higher expected payoff).
I got it to a decent state, but didn’t know how to propagate it or inject it into social communities. I wanted people to be able to tag it on Facebook, and it would reply with an informational card with the analysis and summary.
I feel like machine learning isn't at the level where it can tell if something is misleading, unless it's from a known sketchy source.
Specifics: I'm tracking the edges of the knee meniscus from time lapse video (~ 1000 frames) to measure its deformation under load. This is in the context of research to prevent osteoarthritis. Due to material rotation and irregular geometry, background edges that started off occluded come into and out of view over time. This tends to confuse both machine vision algorithms and human labelers. Because the tracking is for strain measurements, the tracked edge must be the same in all frames; therefore, the tracked edge is be a foreground vs. slightly less foreground edge of the same material (low contrast), not foreground vs. background (much easier). Only few-shot approaches are likely to save time because only ~ 20 specimens are needed to accomplish the immediate objective, and follow-up experiments will probably differ enough to require re-training.
The current plan is to Google "few-shot image segmentation" and try things until something works or the manual labeling effort finishes first, but maybe one of you knows a shortcut. Work is also ongoing to bypass the problem by enhancing edge contrast or using 3D imaging, but machine vision would be the most cost-effective solution.
I'm looking for a more advanced duplicate files finder in Linux, specially one that can handle folders.
Most tool just return a list of duplicate files, but what I need is to also know is if whole folders are duplicate, or subset of others and have it presented in an elegant way t o resolve conflict. ofc everything is in a big folder and everything is a mess, bit of a hoarding collection of files that got accumulated over the years.
I could probably code some stupid script myself, if I ever got the time (spoiler I probably won't) but I have no idea on how to present the result elegantly. So it would be nice if such tool already existed.
For tools which can find differently names files you have for examples: fslint (gui) or fdupes (cli)
It seems like what we're looking for is content-addressable storage [1]. The theory behind it appears to be based on Merkle trees and cryptographic signatures [2].
IPFS already has an implementation [3] of this, and there are other implementations (borgbackup, restic, zpaq) listed in that link and in the content-addressable storage article. Disclaimer: I haven't used any of these yet, just found them a few minutes ago.
[1] https://en.wikipedia.org/wiki/Content-addressable_storage
[2] https://gist.github.com/mafintosh/464bb8f1451f22c9e5c5
[3] https://discuss.ipfs.io/t/ipfs-and-file-deduplication/4674
You can help by trying the API and giving honest feedback.
1. Image generation automation: Like into automation like posting tweets as image to instagram.
2. Image generation scaling: Create multiple images with just variable text, like greeting messages or product images or open graph images.
Some feedback on the Designer:
The size setting dropdown is quite strange. Choices of unfamiliar destinations don't seem to make sense, and some of the things that do seem familiar come out an unexpected size and shape (e.g. "infographic"). The pixel sizes are clearer, but better would be handles on the canvas that can be dragged. There's also a typo: "Choose form a list of sizes".
Circles don't get resized? I can drag handles to make the apparent bounding box bigger, but the circle doesn't change: https://imgur.com/a/KhYluKj. Other shapes seem okay. Chrome 89 on Linux.
The tutorial walkthrough pops up every time I go to the designer, even though I've been right through it.
Fixed most of it.
Your service could provide pre-made templates and an editor, and expose textfields, images, fonts, etc options via URL parameters. Then your service just has to render the SVG and return it as an image with the requested dimensions/format.
Example:
`https://img.bruzu.com?s=<TEMPLATE ID>&title=Hello&font=arial&width=800&height=480&fmt=png`
You could also pass the raw SVG source as a parameter as well, maybe with base 64 encoding or something like that.
I have a plugin that exports some WooCommerce orders into XLS. I would like to add a progress bar via AJAX, because the export may take very long for thousands of orders. But I am not really sure how to use AJAX in context of Wordpress specifically.
I would love to see a minimal functional example, a simple plugin that does something similar. So far, all the plugins I saw were pretty convoluted and I lost my track around the code.
(On a related note: a library of elementary examples for Wordpress plugin development would be nice. Like "This is how you create a menu entry.")
I guess you already have an URL on your Wordpress setup that triggers this export. Let's call it {url}/export.
Wordpress already has jQuery by default included. So'll you need to call that URL using jQuery $.post and then, accordingly to the response, update your progress bar.
There is nothing specifically about Wordpress on this, besides the fact that you need to setup your own URL on Wordpress to do this, and then include your own JS after jQuery. That's all.
If you find this too-complicated, a quick-hack is to create a page on WP Admin called Export Tool, and then on your theme create page-export-tool.php. That .php will be called when visited that Export Tool page.
I am looking for my first client: Ideally someone in charge of a Museum/gallery or other grandiose indoor space.
I wish AirBnBs and hotel rooms would offer this type of preview of their premises.
Edit: There's a bug where if you start dragging with the mouse and let go the mouse button outside the 3D view, it acts like the button is still held down (a bit like mouse capture) which was easier but quite confusing.
It seems unusable with a touch screen on desktop.
The fact that you can fly is useful but non-obvious. I ended up down at floor level and wondering how to see the pictures on the walls.
The demo is in a /fr/ path, but is in English (Chrome offers to translate it to English, because it somehow thinks the English words are French), but then some parts of the interface like "Share your place" are in French.
Additionally, I think it would be best if by default you were stuck to standing head height, then you can either provide buttons to actually move up or down, or lean more into the game aspect and allow the user to jump. Right now it feels like you are floating around with a little drone or something.
In the vein of controls, please please please support WASD too. I understand if you instruct with arrow keys, since for non-gamer users it might be more obvious, but support WASD (or equivalent of what WASD is in QWERTY keyboards) anyways, for 2 reasons: it is much more ergonomic for people who use a mouse on their right hand (I myself am left-handed but use the right hand for mouse anyways), and is more ergonomic for some laptop users, since many laptops have half-size vertical arrow keys which are uncomfortable to press all but momentarily.
As for the head, yes you are right : I should add 'something' that tilts a bit the head up/down when needed.
1. Your demo level suffers from Z-fighting in a few places on the floors and walls. https://en.wikipedia.org/wiki/Z-fighting
2. When viewed full-screen on a 4k monitor, textures are too low resolution. A handwritten note on the wall is unreadable.
3. Lighting is too simple. Because that’s not an FPS shooter you probably don’t need dynamic lightning nor day/night cycle, but it’s still hard. Ideally you need these multiple PBR textures everywhere, and correspondingly complicated pixel shaders.
2: yes but it is a tradeoff between texture quality and minimazing loading time.
3: Yea, ideally. But KISS is my priority : We have an editor that aims to be simple enough for all, thus no BPR & no shader.
3 — I see. Still, you could pre-compute local illumination automatically in the editor, and bake it somewhere. Maybe into vertex attributes, maybe into another lower-resolution R8_UNORM set of textures.
3- --> I will see with client feedback. I do not want at this point to over-engeneer free-visit. First I must find my market.
The preprocessing to increase DPI to 300 did not help when I tried Tesseract, unfortunately. It’s hard to achieve a good contrast between the numbers and the backdrop
If that doesn't work you'll have to add a text localization model to your pipeline.
There will be an appending WAL for batching the writes and for transaction support. A single checkpoint worker will apply the updates from WAL to the index storage periodically. The updates will be batched, to minimize the random seeks.
Please help me evaluate the three indexing approaches, on the criteria of fast read, fast write, cache friendliness and ease of supporting atomic update during checkpoint.
By web search I found this to tutorial to put sentences in an embedding space: https://github.com/BramVanroy/bert-for-inference/blob/master...
I did not read this and am not endorsing it, but it looks like it’s doing roughly what I’m suggesting.
First, differentiating code to build a client side predictor with privacy as a consideration. I have code that describes how to translate domain messages into state changes, and I'm trying to figure out how to predict the effects of sending a message on a client even though the client has imperfect knowledge.
Second, AI for games with imperfect information. Specifically, how to build an AI for Battlestar Galactica.
These are in context of http://www.adama-lang.org/
A secure client side encrypted document (Passport / Id) sharing service
Needs a DSVGO lawyer as part of the team.
Your call, of course. I'm just a random internet stranger spitballing ideas here for how to take that next step and I know next to nothing about this problem space.
Best of luck, whatever you choose to do.
* Make sure you read the rules for Show HN and use your own judgement there as to whether this qualifies.
My problem is, would this be a viable product for companies? It was always a hassle when I worked in labs but I am not sure if it is a big enough problem to design an entire IoT product around.
It took a while but finally came up with something where it was actually closer than I thought.
Thanks for the inspiration.
Solved it.
Does anyone know of a technology which provides similar abilities to EFAB but can also print features with non conductive materials?
Aluminum can be anodized, then sealed. As can a number of other metals.
I want to add IoT tools in several apartments and allow my guests to control them through an app. Currently I already have an app, but no IoT integrations (basically just reservation).
How to do it safely? Also, which automations would be nice to have but not so obvious?
Thinking from things like light control, but also checking the apartment energy consumption in order to detect appliances that need maintenance before they break.
Cut out the sh dependency and just use the "obvious" tech; make the site capable of generating itself, without reliance on any other tools.
Regardless on the answer, I think this gonna be way too slow for your use case. Especially if you want the global optimum. Wikipedia says the current state-of-the-art is randomized algorithms, I don’t think these algorithms are looking for global optima.
I don’t recommend solving that, I think you’ll waste your time. How’s that related to the keyboard anyway?
1. Find all vertices with degree N > 40 (eg: find all points in the graph with more than 40 outgoing edges).
2. For each pair (a, b) of these degree N > 40 vertices, find the set of common points (c) connected to both a and b by edges, eg: there exists 2 edges (a - c) and (b - c). In essence, you're forming V shapes (a - c - b) where the two tips (a, b) of the V have at least N > 40 outgoing edges.
3. Identify the pairs of vertices (a, b) that are connected by 40 or more Vs (eg: there are at least 40 pairs of (a - c) and (b - c) edges).
4. Remove the (a - c) or (b - c) edge with the highest weight. If (a - c) and (b - c) have equal weight, remove the edge which is connected to the node a or b with more outgoing edges.
5. Repeat steps 3 and 4 until each (a, b) pair has at most 39 Vs connecting them together.
At this point, I think you can color the graph with N=40 colors (one color for a and b, then different colors for each in-between point of the 39 Vs between them).
There might be a way to improve the criteria for which edge to remove in step 4 (maybe using the backtracking approach mentioned by other commenters), but this should be a decent starting point.
On another note if the goal is just to avoid files/writing to disk then a bunch of ffmpeg splices to named pipes as inputs to another ffmpeg command to merge them could do the same without the command soup.
https://superuser.com/questions/587511/concatenate-multiple-...
How does one solve for risk in the markets? As in, mathematically. How does one do short term predictions of prices, with a day or two as the prediction range, with a probability > 0.5.
Your idea, if implemented well, may end up being a net positive for society, but I can't help imagining a future where every child, from the moment they are born, has a biometric ID connecting them to a consortium of companies which provide their education, health care, housing, energy, internet connectivity, transport, media access, and so on.
It would be like living in a company town, being paid in company scrip, except you wouldn't notice the restrictions (as long as you kept earning). If you ever increased your income, your consortium might let you choose whether you want to upgrade your housing or your health care plan, but if you lost your job, they'd force you to take one of their choosing and downgrade your plans if it had a lower salary.
In this dystopia, all consumables from food to toilet paper would presumably be sold by Amazon, and other items like furniture and electronics would be provided as a service so that you rent them from your consortium. The only question is why people wouldn't try to undo this system through the political process, but then we might ask that about the current system.
Thank you!
> but I can't help imagining a future where every child, from the moment they are born, has a biometric ID connecting them to a consortium of companies which provide their education, health care, housing, energy, internet connectivity, transport, media access, and so on.
My idea will not have that side-effect because not only is the project non-profit and open source, but it is also decentralised. And if we keep thinkng of a dystopian future then we won't be able to do anything positive unless we become some sort of social revolutionaries. I don't have those skills. But I can think of small ideas to make a positive impact that benefits everyone though. The idea I listed in my original post above is very simple, it helps teachers, lecturers, or anyone who contibutes to a persons monetary decent life via education get appropriately paid for their efforts, and everyone(i.e. businesses) who benefits from an educated person should contribute towards that.
It is a simple idea but notoriously diffcult to deploy because there is a possibility of this getting caught up in a political slugfest.
You can prove two balls are different colors to a colorblind person by having them show you two balls of X color and proving to them that you can differentiate them (watered down example), but it requires validated externalities (ex, you can see colors they can’t).
Defining the external validators is the hard part.
The best option I tried is Google Cloud Vision, but it's still not accurate enough, and it could get quite expensive for large tasks.
Anybody knows a good software for that?
Goal is to bring transparency to medical bills and remove the unknown.
It's been slowing going getting all the maths done.
The average CTO is not as knowledgeable as you might think.
For example, there are CTOs today building applications using no-code platforms, without any academic or technical background. Those people would have much to learn from an software engineering intern at any company.
CTO is a job title, and each company can grant that title at their discretion. Or Technoking, or any other title you might think of.
I am in a community with some CTOs in Latin America and I see this struggle happen every time.
VCs there suck, and do not understand that VCs are about inherently risky investments. They want the guarantees of a low risk business with the profitability of a high risk business.
Leadership in Latin American companies also sucks. As soon as the company has any revenue, the leadership will go all out and spend it all on themselves on some extravagant lifestyle adjustments rather than reinvesting it on the company. And because of this, many companies stay small and mediocre without fulfilling their true potential. They waste their money on MBAs learning things they are unwilling to apply.
Then, there is nepotism. Hiring from your family equates creating conflicts of interest, creating situations where relatives of key people do not need to comply with HR, do not need to be competent, cannot be fired and create an unprofessional atmosphere.
Compensation in most Latin American companies sucks. They aspire to be like Silicon Valley startups and hang framed posters of Steve Jobs on their wall, but when it's moment to create compensation packages they grant zero stock and award zero bonuses, all while having the same work-life balance of a startup. They do not understand the key role of employee stock in company growth.
Why the fuck would you work for a wannabe startup that doesn't give you any stock or bonuses? Or work for some entitled aggrandized clown that doesn't understand the simplest technical concepts? And the answer is: because business owners in Latin America have never had to care much about their workforce. Because of exploitation and their informal caste system, having a happy workforce has never been a requirement.
That's why Latin America is undergoing a massive brain drain that will only get worse in the years to come.
Yeah, Latin America sucks, I already know that.
Now back to my problem...
I just accepted an EM role at a FAANG. This is a career "boomerang" for me - I was an engineer in the past, but then moved into technical support management. I'm coming back to engineering, but this is the first time I will have done it at a large company. All of my engineering experience has been at small scrappy startups where we just did everything, did it fast, and prioritized by whatever was most on fire. I don't think I've ever actually done a proper "sprint". While I have written a lot of code, the workstyle on my new team will be almost entirely foreign to me.
Whose got pro tips for leading an engineering team in a large organization? What makes a high-powered team? What are the easy mistakes that will drive us into the ground?
This guy made the switch from dev to EM and has written articles on running effective teams,(e.g. one on how to help devs act as project leads) and even has resources like the docs he sends to tech leads when they start a new project, w/ lists of responsibilities etc: https://blog.pragmaticengineer.com/things-ive-learned-transi... (just linking to the most directly applicable article but definitely browse around)
Also, always avoid meetings to ideate. These are some of the most common and they are a huge waste of time compared to listing out ideas in a doc and having people review that asynchronously. And yet these meetings have a tendency to get called all the time. For example "there was a fire, let's get all the leads/EMs/directors together for 1 hour to figure out how to avoid this next time".
Your two most important duties are 1) making sure your developers are given space to implement what's important and 2) building relationships throughout the company to better anticipate future needs that helps you do #1.
Happy to chat anytime as well, email in profile.
I'm curious on how other people solved it ( by cookies, subdomain, ... ) and if you used a JwtToken for it.
But I haven't decided yet on the actually flow. Where I'd identify the current tenant or impersonate him.
+ The influence of impersonation on that flow.
Deciding which features to ship and what to skip is a bear — it's a big product design space
Long version: My Mu project (https://github.com/akkartik/mu) is building a computing stack up from machine code. The goal is partly to figure out why people don't do real prototyping more often. We all know we should throw the first one away (https://wiki.c2.com/?PlanToThrowOneAway), but we rarely do so. The hypothesis is that we'd be better about throwing the first one away if rewriting was less risky. By the time the prototyping phase ends, a prototype often tacitly encodes lots of information that is risky to rewrite.
To falsify this hypothesis, I want to make it super easy to turn any manual run into a reproducible automated test. If all the tacit knowledge had tests (stuff you naturally did as you built features), rewriting would become a non-risky activity, even if it still requires some effort.Turning manual tests into automated ones requires carefully tracking dependencies and outlawing side-effects. For example, in Mu functions that modify the screen always take a screen object. That way I can start out with a manual test on the real screen, and easily swap in a fake screen to automate the test. Hence my problem:
How do you represent screens in a test?
Currently I represent screens as 2D arrays of characters. That is doable a lot of the time, but complicates many scenarios:
* Text mode character attributes. If I want to check the foreground or background color, I currently use primitives like `check-screen-in-bg` which ignores spaces in a 2D array, but checks that non-spaces match the given background attribute. In practice this leads frequently to tests that first check the character content on a screen and then perform more passes to check colors and other attributes.
* Non-text. Checking pixels scales poorly, either line at a time or pixel at a time. A good test should seem self-evident based on the name, but drawing ASCII art where each character is a pixel results in really long lines or stanzas. So far I maintain separate buffers for text vs pixels, so that at least text continues to test easily.
* Proportional fonts. Treating the screen as a grid of characters works only when each character is the same width. If widths differ I end up having to go back to treating characters as collections of pixels. So Mu currently doesn't support arbitrary proportional fonts.
* Unicode. Mu currently uses a single font, GNU Unifont (http://unifoundry.com/unifont/index.html). Unifont is mostly fixed-width, but lots of graphemes (e.g. Chinese, Japanese, Korean, Indian) require double-width to render. That takes us back to the problems of proportional fonts. Currently I permit just variable width in multiples of some underlying grid resolution, but it feels hacky.
Can people think of solutions to any of these bullets in a text-based language? Or a more powerful non-text representation?
I'd consider taking inspiration from the following sources:
1. GUI toolkits like QT QML [1] or Android [2]. These typically build a hierarchical tree of different components (eg: start with a root window, which contains panes, which in turn contain text and buttons). Each component may contain different properties (eg: font, color), and properties may be inherited from the parent component.
Advantages:
+ preserves semantics of component properties and how they are linked to each other (eg: the caption is below the image)
Disadvantages:
- complexity: building a layout/constraint engine can be difficult, or alternatively you can use absolute positioning with relative offsets which can be tedious to use (in this case the layer-based approach below might make more sense).
[1] https://en.wikipedia.org/wiki/QML
[2] https://developer.android.com/guide/topics/ui/declaring-layo...
2. Graphical editor programs like Gimp or Photoshop, or Adobe Flash.
These build up a screen as a collection of vertically stacked layers or assets (eg: graphics, text, etc) with attached properties and optionally bounding boxes. Higher layers/assets occlude the content of the layers below them, so you would need to implement some kind of visibility logic.
Advantages:
+ simplicity
+ you can use identifiers for assets, and therefore don't need to perform pixel-by-pixel comparisons.
Disadvantages:
- may lose some information about how different components are related to each other
Also, rather than a raster pixel-based representation, it might make sense to use a vector representation internally [3]. The most popular vector representation is SVG. The full spec is very verbose, so you probably only want to implement a small subset of it. This would permit you to specify properties like line thickness, color, striped/dotted patterns. At render time, you could convert the (proportional) fonts to vectors as well for consistency, and then rasterize the entire scene when rendering to a display surface. But for testing, it would be better to use the scene graph / vector format which is easier for users to reason about.
[3] https://en.wikipedia.org/wiki/Vector_graphics
[4] http://blog.leahhanson.us/post/recursecenter2016/haiku_icons...
But perhaps this is over-complicating things.
Android includes the Espresso UI testing framework [1]. Essentially, you can specify matchers that compare your expected values or predicates against an actual object identified by an R.id identifier. It's very powerful (since you can write your own custom matchers) but can be cumbersome to use [2].
[1] https://developer.android.com/training/testing/espresso/basi...
[2] Example Espresso Test: https://github.com/android/testing-samples/blob/main/ui/espr...
https://github.com/android/testing-samples
Alternatively, Squish [3] is a very polished and more elegant commercial testing tool that lets you record test-cases using a GUI tool and convert them into (ideally modularized) methods that verify object properties or compare (masked) screenshots of the GUI:
[3] https://www.froglogic.com/squish/features/
Demo video (starting at 14:24): https://youtu.be/ElH-3MVHPRw?t=864
They abstract away a lot of the functionality using the Gherkin [4] domain-specific language so that tests are easier to read at a high level (but you can still dig down into the underlying programmatic implementation).
[4] https://cucumber.io/docs/guides/overview/
This is probably too much complexity for your use-case, but may provide some ideas or inspiration for what is possible. Perhaps a simplified matcher-style system might be a good starting point though.
Low code for devs. https://github.com/hofstadter-io/hof
Trying to reduce redundant tasks and simplify changes with minimal effort.
When proof-of-stake takes over, there won't be any miners. The block proposal process is done by stakers instead. Some of the incentive issues with miners still exist with stakers, but raw competitive power consumption isn't one of them.
It's true that proof-of-stake has been talked about for years, but it has picked up momentum since late last year, as the staking network was actually launched.
The proof-of-stake network has been staking real ETH since end of last year, but does not yet handle mainnet ETH contract transactions. It's called Eth2, but that's caused some confusion, because it's not really a second version to run alongside the first, it is the R&D branch into proof-of-stake and other technical improvements, with mainnet ETH expected to adopt it in due course.
So, the Eth1 components have been renamed "execution layer", Eth2 components renamed "consensus layer", and through a series of testnets and API developments which have been quite active this year, a big change called "The Merge" is being worked on by multiple funded groups (for client diversity) of core Eth developers at the moment.
The Eth2 staking network that already exists has demonstrated the viability, and the investment of real ETH in serious quantities has built up some cryptoeconomic stability prior to its deployment as the ETH consensus layer. The time lag is intentional - you don't want to suddenly switch all ETH over to a network with too few invested stakers.
You originally said you want to destroy proof-of-work blockchains, and were looking for advice on how to do that with ETH. It's already being destroyed on ETH by proof-of-stake. You asked for advice, that's the advice. You don't need to do anything except wait.
Now you are saying you don't need to put effort into destroying proof-of-stake. Why is that relevant here? It suggests to me your goal is different from what you originally stated. Are you looking to see the destruction of more than just proof-of-work? The destruction of BTC and ETH, even if they switch away to another consensus mechanism?
Plenty of problems - none of them technical - all people problems!
1. I want to create way to generate electrical power without pollution. Basically, a closed cycle process that releases no pollutants, or electronic waste.
2. I want to do everything I can to eliminate gender bias in the world.
We have a small hydro-electrical plant una River near my house and really it's no big deal, it fits very nicely in the surrounding environment and it produces clean energy.
It's also educational because since the river is near the city small children classes can visit it and learn about it.
Geothermal can be a solution for generating electricity directly, but if you'd like to minimize electronic waste perhaps it would be easier to use it to replace alternative energy sources for HVAC purposes.
Biofuels (eg: plant bamboo, grow it, then burn it) can also technically be closed cycle energy sources.
Solar water heaters can also reduce electrical or fossil-fuel-based energy consumed for generating hot water.
What progress have you made in your work on either front? What sort of work do you do to solve these problems?
I am also spending much time lately in the SF kink community to build a fundamental understanding of the biases people have experienced in life with respect to their gender identity, and am strongly considering HRT so I can live life on the other side and experience the prejudice first hand.