Added an entry for my data visualisation tool here: https://github.com/stared/interactive-machine-learning-list/....
Edit: found an updated link for seeing theory so I fixed it in the PR above. Feel free to cherry-pick if #24 is not relevant.
107 karma · joined June 18, 2021
Added an entry for my data visualisation tool here: https://github.com/stared/interactive-machine-learning-list/....
Edit: found an updated link for seeing theory so I fixed it in the PR above. Feel free to cherry-pick if #24 is not relevant.
To give a sense of what the loss value means, maybe you can add a small explainer section as a question and add this explanation from Karpathy’s blog:
> Over 1,000 steps the loss decreases from around 3.3 (random guessing among 27 tokens: −log(1/27)≈3.3) down to around 2.37.
to reiterate that the model is being trained to predict the next token out of 27 possible tokens and is now doing better than the baseline of random guess.
This resonates with my startup experience, especially when the organisation is growing rapidly (which is almost always the case for VC-funded startups)
> in a way it can be more honest for a developer with any commercial intentions to not open source something while offering paid ways to modify the source code
I appreciate that you made your commercial intentions clear in the comments and I agree that you don't want to walk back your license as it definitely leads to backlash. I think developers (i.e. people on HN) tend to trust open-core models more where there's a clear separation between the core open-source components and a paid offering but I guess you don't have to overindex on those given the general audience of this tool.
I don't know what your ultimate monetization goal is, but in my opinion one realistic path is an one-time only payment of some reasonable amount with the full feature as opposed to subscription (I saw your Stripe link in the source).
> Scout was our most-used product. Its users weren’t our target market, but some were. We had vague ambitions to use it as an inbound marketing tool, but we never capitalised on it. This was a missed opportunity.
was applicable in a conservative industry like real estate and government as lots of open-core companies operate on this model (e.g. free open-source software and marked-up hosted solutions).
This reminded me of the summer when I accidentally landed my first programming gig at my friend's startup as a part-time web developer during university. I was running out of money for food as my research assistant role was unpaid and I asked around for some odd jobs in the city. As a non-CS major, all I had was data analysis skills in Python but my friend thought I was was good enough to help him out with his SaaS startup, and I ended up learning a lot on software engineering principles on that job.
Also thanks for the error checker! I pushed the fixes in https://github.com/visprex/visprex.github.io/pull/4
Yes I had the same experience for analytics work some years ago. As others have pointed out, Visprex only works in a happy path where data is a clean CSV file so will definitely need to work on data cleaning. I have a DuckDB integration planned but not sure if this is easy enough for the target audience. Will try to add some predefinied functionalities, thanks for the feedback!
> With the UI I want to be able to toggle between different strategies quickly - strip characters from a column to treat it as numeric, if less than 2% or 5% of values have a character, fill na with mean, interpret dates in different formats - drop if the date doesn't parse
Those are really good examples and I can make those predefined preproccesing techniques available as toggles in the dataset tab. Thanks for the feedback!
I would say the data loading functionality compares very poorly to Excel CSV import for all the reasons you pointed out, and I agree that the users can face those formatting issues which could be resolved in another tool like Excel or Google Spreadsheet for non-technical users and Notepad++ or editors for a bit more technical users. The assumption on CSV files being clean is strong so I will try to surface import errors at least, and in the meantime point to different ways to format the data as those tools will be complementary to Visprex.
> To me, this implies that no steps have been taken to manage user/data privacy.
This is a good point. I fixed the wording and now it simply reads "No tracking or analytics software is used". Thanks!
Some of those plans are mentioned in my blog post reflecting on building this app: https://kengoa.github.io/software/2024/11/03/small-software....
Do you have the results of test262_runner.rb? I came to know about test262 at a talk by the porffor's author and something like https://github.com/CanadaHonk/porffor?tab=readme-ov-file#tes... in README would be great to show this progress. Great project by the way!
This hits so close to home. The difficult part is that this tendency is not correlated with the level of urgency of the topic or priority within the project.