HNHacker News
TopNewBestAskShowJobs

vikp

1,176 karma · joined August 13, 2012

I used to teach people, now I teach machines.

Email me at hn@vikas.sh, or check out my work at https://www.vikas.sh.

submissionscomments
vikp··on StableLM: A new open-source language model
It's unclear which models will be trained to 1.5T tokens. The details of how many tokens each model saw in training are on Github - https://github.com/stability-AI/stableLM/ . But only for the ones that have been released.
vikp··on StableLM: A new open-source language model
It's fantastic that more orgs are releasing open-source models trained on more than 300B or so tokens. Here's my take from the details I could find.

Pros

  - 4096 context width (vs 2048 for llama, gpt-j, etc)
  - 3B to 65B released or in progress
  - RL tuned models available
  - Trained on more tokens than existing non-llama models
  - 128 head dim, so can use flash attention (unlike GPT-J)
Cons

  - No benchmarks released, or details about the model
  - Somewhat restrictive license on the base models, and NC license on the RL models
  - Small models only trained on 800B tokens, compared to 1T for llama-7B, and potentially more for other upcoming alternatives (RedPajama, etc).  I'd like to see their loss curves to see why they chose 800B.
High-level, this is likely to be more accurate than existing non-llama open source models. It's hard to say without benchmarks (but benchmarks have been gamed by training on benchmark data, so really it's just hard to say).

Some upcoming models in the next few weeks may be more accurate than this, and have less restrictive licenses. But this is a really good option nonetheless.

vikp··on Ask HN: I am a full stack developer, where do I start learning for AI
I've been making a course to help with this - https://github.com/VikParuchuri/zero_to_gpt . Lots of coding, with optional videos. It has very few prereqs.
vikp··on How can I prevent my site from being a free dataset for LLMs?
Attribution comes to mind.
vikp··on Show HN: Tabby – A self-hosted GitHub Copilot
How does this compare to fauxpilot - https://github.com/fauxpilot/fauxpilot? Fauxpilot also uses Triton with fastertransformers and GPT-J style models (codegen).
vikp··on EleutherAI announces it has become a non-profit
If you want more open research and weights, you should be happy about this announcement. Incorporating as a nonprofit doesn't guarantee that an organization will act ethically, but it does make it more likely. Nonprofits have more restrictions around how they can spend their money. They also pay less, so the teams is more likely to be mission-aligned.

The profit motive pushes organizations to keep their work secret (which is what has happened with OpenAI).

vikp··on Microsoft is preparing to add ChatGPT to Bing
In this particular case (summarizing search results), it seems to work well. I think anchoring to known information helps avoid a lot of the issues (hallucinations, boring language, etc).

You can check out an open source demo I made if you want to play around with how search + GPT work - https://github.com/VikParuchuri/researcher .

vikp··on Microsoft is preparing to add ChatGPT to Bing
If you want to try this out today, I made an open source version using Google + GPT-3 - https://github.com/VikParuchuri/researcher.

It works by getting Google results, finding the most relevant text chunks in the pages, then passing them to GPT-3 to generate a summary. The summary includes sources, so you can verify the info.

It works well for research questions like "what is the best smartphone with a small screen?", or "how do large language models work?". It gives better answers than ChatGPT alone for these types of questions.

I'm finding that the biggest time savings is from not having to wade through SEO-spam pages to find the one relevant paragraph.

vikp··on Show HN: Researcher – answer questions using Google and GPT-3
Hi HN - it's been getting hard for me to do research with Google. If I'm looking for the best smartphone, or the right Javascript framework, I have to wade through dozens of SEO-spam pages to find the answer.

I've experimented with different solutions to this - Kagi, DDG, SearX, and writing my own custom filters. Nothing quite works the way I want.

Last weekend, I decided to feed text from Google search results into GPT-3 to generate summaries with citations. This works well - I get an overview of the topic, but also dive deeper if I need to by clicking on the citations. Citations help verify the information and ensure accuracy (I've noticed fewer hallucinations than with GPT-3 alone).

It works by getting results from Google, scraping the pages, extracting text chunks that align with the question (using embeddings), then sending the most aligned chunks to GPT-3. The GPT-3 prompt specifies to only use the chunks to generate the answer (this usually works).

Since it's self-hostable, it's easy to tune and hack to suit your preferences.

vikp··on ChatGPT is a ‘code red’ for Google’s search business
I've had the exact same problem, and have experimented with Kagi, DDG, etc, to try to find better results.

ChatGPT, in my opinion, is great for "how do I code X" type questions, but isn't so good at the types of queries you mentioned, due to the lack of a search engine.

My weekend project was an open source combination of Google + GPT that returns pretty good results for these types of queries. You can check it out here - https://github.com/VikParuchuri/researcher

Example - the response to "what are the best current smartphones" is:

`...According to Search Result [2], the best phones have been thoroughly reviewed and tested, and include the Apple iPhone 14 and 14 Pro, the Pixel 7 Pro and the Samsung Galaxy S22 Ultra. Search Result [5] also states that there are strong options available at all price levels, so you don't have to spend a lot to get something great...`

vikp··on ChatGPT is a ‘code red’ for Google’s search business
It's doing abstractive summarization over the search results, using GPT-3. The pipeline is:

  - Search using Google
  - Run some filters to exclude SEO spam, etc.
  - Scrape the pages that are returned
  - Find chunks of text likely to align with the answer (comparing embeddings)
  - Feed the most likely chunks into GPT-3 to get a summary
It is leveraging GPT-3 to produce better summaries, and it isn't purely extractive - the LLM uses context and knowledge to generate a better summary.

I want to experiment with a local model next, versus using GPT-3.

vikp··on A new chat feature has been released by You Search
This is very cool! It's nice to see other implementations outside of ChatGPT.

I noticed that it doesn't always cite sources, so I assume that some results are coming straight from an LLM, and some are web search + LLM.

If you're always looking for sources to be cited, I built a project this weekend that will summarize a topic and cite sources - https://github.com/VikParuchuri/researcher . It's definitely slower than this implementation, but it is self-hosted and hackable.

vikp··on ChatGPT is a ‘code red’ for Google’s search business
I don't think it's a huge lift to restrict a language model to "known" good facts from search results. And to have it cite sources.

I made a proof of concept this weekend - https://github.com/VikParuchuri/researcher . There are some issues, but it's very useful.

vikp··on How can I find someone to explain, say, Git to a non-technical person
I built a side project recently to help with this. It searches Google, then feeds the relevant results from the pages into GPT-3 to get a summary. It seems to be accurate so far in my testing - https://github.com/VikParuchuri/researcher .
vikp··on I am done. I give up
Don't start a business unless it's a pain point you are personally invested in. Lifestyle-wise, it's a grind with mental health challenges along the way, as you pointed out. Salary-wise, it's unlikely that your own business will ever match a regular developer job. It took me 4 years of bootstrapping before I had a decent baseline salary, and it's still lower than my ML engineer salary 10 years ago.

I don't get the current culture of starting a business for the sake of starting a business. The lifestyle can be great after a few years when everything is running smoothly. But the marketing ignores the fact that it takes multiple years of grinding to get there, with no guarantees.

Specific to your situation, it sounds like you were reaching the wrong audience. You need to reach the audience interested in your product, not an audience of other entrepreneurs. This takes a lot of time and energy. Marketing needs to be more than 50% of your effort these days. Content marketing on social, blogs, and Youtube tends to do well, but takes a couple of years to ramp up.

Thanks for sharing, and I'm glad you made the right decision for you. Mental health is the most important asset you have - you're wise to protect it.

vikp··on Ask HN: Can you help remember the name of a project shared on Hacker News?
I haven't seen it specifically on HN, but SearX does what you described - https://searx.github.io/searx/ .
vikp··on Riffusion – Stable Diffusion fine-tuned to generate music
Producing images of spectrograms is a genius idea. Great implementation!

A couple of ideas that come to mind:

- I wonder if you could separate the audio tracks of each instrument, generate separately, and then combine them. This could give more control over the generation. Alignment might be tough, though.

- If you could at least separate vocals and instrumentals, you could train a separate model for vocals (LLM for text, then text to speech, maybe). The current implementation doesn't seem to handle vocals as well as TTS models.

vikp··on How does GPT obtain its ability? Tracing emergent abilities of language models
Large language models like GPT work by generating the probabilities for the next word in a sequence, given the previous words.

You can make this purely deterministic (same sequence every time) by just selecting the word with the highest probability repeatedly.

Most models will inject some randomness by sampling from the probabilities, to make the generated text more realistic. I don't know what parameters they use for ChatGPT, but they are likely injecting some randomness.

It's also likely that they are continuously training ChatGPT with reinforcement learning, so responses may change over time due to this, also.

vikp··on How does GPT obtain its ability? Tracing emergent abilities of language models
It's unclear to me how you could separate knowledge and reasoning:

- Reasoning typically requires base knowledge to work from. A side effect of training reasoning is embedding knowledge into the model parameters.

- Even if you offload the search portion (either through outputting special tokens that are postprocessed, or applying the model in multiple steps with postprocessing), you still need embedded knowledge for the model to decide what to search for, and then to successfully integrate that knowledge (in the multi-step case).

Maybe some kind of post-facto pruning of model weights?

vikp··on Solving brain dynamics gives rise to flexible machine learning models
I was curious to see an implementation, and I found this code for an earlier version of CFCs - https://github.com/raminmh/CfC .
vikp··on Ask HN: Best sleep trackers?
I haven't seen Withings Sleep mentioned - https://www.withings.com/us/en/sleep .

This one is basically "set it and forget it". You put it under your mattress and calibrate it, then it tracks automatically.

I like it because I don't have to bring any gadgets into the bedroom, or remember to turn on tracking.

It seems pretty accurate from what I've seen so far, although I don't have anything to compare it to. I've found it very useful for testing various sleep interventions.

vikp··on Lambda School (YC S17) now pays eligible students $2k/month
If you're only getting repaid when someone is employed and making over 50k, which seems to be the case from your website ("Upon completion of the program, students will pay 10% of their [pre-tax] salary for a five year period once they're making at least $50,000 per year"), then the "repayment floor" for Lambda school is 25k (with probably a close-to-50k "repayment average").

I'm assuming that if someone loses a job, and later regains one, then payments pause while they're unemployed, but don't count for "repayment time", due to an article I read -- https://www.nytimes.com/2019/01/08/business/dealbook/educati....

How I think it works is that once someone graduates, they have 5 years of "repayment time debt". Once they get a job that pays >50k, the "repayment time debt" starts ticking down. If they lose their job, or start making <50k, then their "repayment time debt" stops ticking down, and will resume when they next get a job.

Even with a 10-year cap, this seems very close to traditional debt. It seems likely that software engineers will make more than 50k for 5 out of 10 years.

vikp··on Lambda School (YC S17) now pays eligible students $2k/month
This is very interesting, and I'm curious about the nature of the risk that you're taking on.

It seems like risk would only become a factor for Lambda School if you're taking 10% of pre-tax income from graduates for the 5 chronological years immediately after they graduate (regardless of employment status).

If you're instead taking 10% of pre-tax income from graduates for the next 5 years in which they're working (non-chronological), then the risk on your side actually seems fairly low, since most software jobs in the US pay >50k a year.

If you're taking 10% of pre-tax income from folks for the next 5 years in which they're employed AND making >50k, then the risk for Lambda School is basically zero. I believe this scenario is what is actually happening.

If the final scenario I laid out is what is happening, then this seems very close to traditional debt (take a loan out for living expenses + tuition, then repay a much larger amount down the line).

vikp··on Ask HN: Do you know a less distracting Slack alternative?
We moved to Twist (https://twistapp.com) a few months ago after having similar issues with Slack. Twist is more forum-like, so you avoid the "I have to jump in now" feeling that a continuous chat stream gives you.
vikp··on Ask HN: Who is hiring? (February 2018)
Dataquest | SF | Director of Marketing | Remote, Full-time | $90k - $130k | https://www.dataquest.io

At Dataquest, we teach data science interactively online to hundreds of thousands of students worldwide. We're focused on teaching skills and building intuition from the ground up with our project based curriculum. Unlike most educational options, we focus on motivating students to learn, not just content delivery. We have students go from no programming knowledge to jobs at companies like SpaceX, Amazon, and Microsoft, and you can read their stories here -- https://www.dataquest.io/stories .

Help us build awareness of Dataquest by scaling our marketing team, analyzing metrics, and running experiments with new growth channels. This is a chance to help students around the world learn while having a lot of ownership over the direction and messaging of the company. Ideally, you'll have experience with running and optimizing ads, analyzing data, and optimizing conversion funnels.

We're a bootstrapped team of 11, and we've been growing 2-3x per year since we launched in 2015, primarily through content marketing. We're looking for someone who can meaningfully increase this growth rate.

If you're burnt out doing work that doesn't feel like it has a direct impact, you have a passion for data science, or you want to peek inside a profitable bootstrapped company, this role could be a good fit.

Please email vik@dataquest.io if you're interested.

vikp··on Steve Wozniak announces tech education platform Woz U
Although it would be unfortunate if this was the case, this paragraph from the about page (https://woz-u.com/about/) leads me to think that this is just using Steve Wozniak's name for branding:

Inspired by Steve Wozniak, co-founder of Apple Computer, we specialize in technology and career-based programs designed to get people into the workforce quickly and affordably...Led by higher education experts, Exeter Education, students will learn the skills necessary to take flight within the technology industry.

It looks like Woz U is affiliated with Exeter Education and Southern Careers Institute. Exeter Education appears to be a new company in Arizona (more info at http://www.azcentral.com/story/money/business/tech/2017/10/1... and https://www.bizjournals.com/phoenix/news/2017/04/14/former-g...).

Southern Careers Institute (http://www.scitexas.edu/) seems to be a vocational school of sorts. Neither of these are bad things, but they temper the initial excitement I had around "Steve Wozniak is launching an online education platform."

vikp··on Ask HN: Sell my startup for $14M because I can't raise $2M?
Not typical, but possible with the right team/tech/connections. OP said that the local VC scene asks for 1M+ ARR to invest, which led me to think that revenue was lower.
vikp··on Ask HN: Sell my startup for $14M because I can't raise $2M?
I haven't seen this advice in this thread yet, but there is a path other than selling or raising VC. It's focusing and growing your revenue until you get through this crunch.

We've been in a similar situation before, that we managed to grow our way out of. It is possible to get through this, keep growing your business, and do it all on your own terms.

To resolve the short-term feeling of being overwhelmed, here are a few ideas:

* Try to get more $$ upfront by converting customers to annual plans, or raising implementation fees.

* Defer some customers until your team is larger (or they pay more).

* List out what you're working on, and ruthlessly trim anything non-essential.

* Contract out what you can.

* Raise some angel or friends and family $$.

I don't know what your revenue is, but I'd guess a few hundred k ARR. In the medium term, if you can grow 5-10% a month, you'll double or triple your revenue in a year. This will give you a lot more optionality in the long term:

* You can just keep bootstrapping forever if you want.

* If you raise, you'll get much better terms with more revenue.

* If you sell, you'll get a higher price.

Of course, this depends on having solid growth channels, and a reasonably sized market.

Taking the time out now to chase VC funding when it's not there will hurt your ability to grow in the short and medium term. If you focus on growing revenue instead, you'll increase your chances at VC funding or a sale in the long term.

I'm happy to chat more if you want -- email is in my profile.

vikp··on Ask HN: I've lost the ability to concentrate. How can I fix this?
What really helped me was to:

* Uninstall some distracting apps, and block the rest with Appblock

* Not check email on weekends, and do outdoor things like hiking instead

* Start running as much as possible, for at least 30 minutes 3 times a week

* Meditate for 10 minutes a day

* Cut out sugar and alcohol

* Install Stayfocusd to limit time spent on sites like HN

* Go to bed by 10 every day, and have 30 minutes of "unwind" time in bed without devices

I don't always adhere to all of these well, but when I do, I feel amazing, and can concentrate extremely well.

I would encourage you to think of this as less of a "quick patch to get back to concentrating on things" and more as "I need to change my lifestyle to improve my mental state. It's going to be a long process, but the long-term rewards are worth it".

vikp··on Ask HN: Who is hiring? (August 2017)
Dataquest | San Francisco, CA | Remote or onsite | Full-time

At Dataquest, we teach data science interactively online to hundreds of thousands of students worldwide. We're focused on teaching skills and building intuition from the ground up with our project based curriculum. Unlike most educational options, we focus on motivating students to learn, not just content delivery.

We have students go from no programming knowledge to jobs at companies like SpaceX, Amazon, and Microsoft, and you can read their stories here -- https://www.dataquest.io/stories . Best of all, we do it at a low monthly cost of $29 or $49.

We're a bootstrapped company, which we think is extremely important, since it aligns our incentives with our students.

Our open roles are a great opportunity if you're passionate about teaching, if you're burnt out doing work that doesn't feel like it has a direct impact, or if you want to peek inside a profitable bootstrapped company.

Please email vik@dataquest.io if you're interested.

Open roles:

* Data Science Instructor (75k-105k) -- outline our curriculum, teach concepts, and analyze data to continuously improve how students learn. This is a chance to make a direct impact on students around the world, while teaching and learning interesting concepts (neural nets, data pipelines, etc).

* Data Analysis Instructor (70k-95k) -- teach data science concepts to non-technical people. Ideally, you have experience with SQL and Excel, exposure to teaching, and, most importantly, are passionate about building intuition by crafting good explanations.

* Devops (75k-100k) -- maintain and enhance our backend infrastructure that allows students to run code and have it automatically checked for correctness. Develop our deployment infrastructure and make architecture decisions. Work with Python 3, Docker, and Kubernetes.

* Data Journalist (60k-80k) -- Help us teach data science on our blog (dataquest.io/blog) that gets hundreds of thousands of monthly uniques. Work with technical experts to build compelling content that has value to students learning data science.

← PreviousPage 3 of 6Next →