HNHacker News
TopNewBestAskShowJobs

tdba

768 karma · joined November 12, 2015

Reach out: getprojectfalcon _at_ proton _dot_ me
submissionscomments
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
It's true that inference is still very often done on CPU, or even on microcontrollers. In our view, this is in large part because many applications lack good options for inference accelerator hardware. This is what we aim to change!
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
There are four big categories of ML accelerators. You already familiar with CPUs and GPUs; then there are FPGAs, which offer better performance and efficiency while remaining flexible. Finally there are ASICs (of which the TPU is an example), which offer the best performance and efficiency but retain very little flexibility, meaning if your ML model doesn't work well on an ASIC then your only option is to change your model.

We chose to focus on FPGAs first because with them we can maximize the usefulness of Tensil's flexibility. For example, if you want to change your Tensil architecture, you just re-run the tools and reprogram the FPGA. This wouldn't be possible with an ASIC. That said, we'll be looking for opportunities to offer an ASIC version of our flow so that we can bring that option online for more users.

tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Awesome, I think your use case would make a lot of for Tensil. Looking forward to chatting more!

The core technology will always be free and open source, so to commercialize Tensil we're planning to offer a "pro" version which would operate under a paid license and provide features specifically needed by enterprise users. We're also working on a web service that will let you run Tensil's tools in a hosted fashion, with one major feature being the ability to search across a large number of potential architectures to find the best FPGA for your needs. Extra paid support and other services will also be in the mix.

tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
I replied to your other comment here about commercialization https://news.ycombinator.com/item?id=30652150 I hope it was helpful!
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Yes, Pynq Z2 should work just as well (it's the exact same FPGA, just a slightly different board). We've been testing with Z1 which is why I recommended it in the tutorial.

For commercialization, the core technology will always be free and open source, but we plan to offer a “pro” version with extra enterprise features under a dual license arrangement, similar to Gitlab. We are also working on a cloud service for running our tools in a hosted setup, in which you’ll be able to run a search across all possible Tensil architectures to automatically find the best FPGA for your model. I'd love to hear your feedback on these plans!

tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Sure, we'll check it out!
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Wow, awesome project! This is exactly the kind of thing we had in mind when we built Tensil. I'd be very curious to hear what happens if you make a v2 perhaps using Tensil for comparison.
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Definitely! You can see some of our benchmarks here, and we'll be expanding this list soon https://www.tensil.ai/docs/reference/benchmarks/
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
You're very welcome! Stay in touch - I've listed some contact methods here and there in the thread, and we'd love to hear from you again.
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Will do - as I mentioned in another comment, it can be a bit subtle to find an apples-to-apples comparison, but we'll soon add some cross-platform that we think are reasonable.
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Thank you - we'd love to see more OSS support from FPGA vendors too and we'll be watching closely for any developments there.
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Thanks, we'll take a look!
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Yep, this is something we've heard before. If you're really familiar with the Xilinx ecosystem, one way we've described Tensil is that it is the "Microblaze for ML" - easy to use, lots of flexibility and customizability, with performance good enough for most applications. The DPU and FINN would then be the more specialized tool for situations where you need specific features they are optimized for.
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Absolutely, the UX for compiler tools often leaves a lot to be desired. This is something we want to fix!
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Generally the comparison between Tensil and any fixed ASIC is going to run along similar lines, which we explain in this comment regarding the Coral accelerator: https://news.ycombinator.com/item?id=30643520#30645318

The big difference is that while those fixed ASICs offer great performance on the set of models they were optimized for, there can be big limitations on their ability to implement other more custom models efficiently. Tensil offers the flexibility to solve that problem.

tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Glad this helped clarify things for you! The tricky thing about benchmarks is that one of the key benefits of Tensil is the flexibility to find a trade-off between performance, accuracy, cost and power usage that works for you. Benchmarks that only consider performance or performance per watt can be a bit narrow from that point of view. That said, this is a good idea and we'll add some comparisons that we think make sense to the docs!
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
This is a great idea, we're looking at boards that could be used in combination with a Raspberry Pi. The reason we haven't investigated this so far is that most of the dev boards we've tested with have an ARM core embedded in the FPGA fabric, so the additional CPU the Raspberry Pi would provide wasn't necessary.
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Agreed, this is the way things seem to be trending. We'll definitely add support for transformers in the near future, the question is only whether there are other things we should work on first, especially with respect to the edge and embedded domain where smaller conv models still dominate. Thank you for the links!
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Coral is a great project, especially if you are using a completely vanilla off-the-shelf model. However if you've ever tried compiling a custom ML model for it, you know how finicky it can be. There are lots of ways that you can accidentally make it impossible for Coral to run your model, and it can be difficult to figure out what went wrong.

With Tensil, you circumvent that problem by changing the hardware to make it work for your model. If you have made modifications to an off-the-shelf model or have trained your own one from scratch, it might be a better option from the point of view of ease-of-use and even performance.

tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
In our current demos, the Tensil logic talks to the host through a couple of AXI and AXI Stream interfaces. There are AXI adapters for many other protocols, including PCIe, that should be able to support many different kinds of connectivity. Here's a link to our docs explaining the host<->Tensil connection: https://www.tensil.ai/docs/howto/integrate/#2-connect-the-ax...
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
We're working on our roadmap right now and prioritizing support based on user interest. If there's a particular model or set of models you're interested in accelerating, I'd love to hear about it!

If there's a lot of interest in transformers, we'd aim to offer support in the next couple of months.

tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
VMAccel looks very interesting! Send me an email and we can explore how to collaborate.
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Thanks! The general answer is that it depends on your model and on which FPGA platform we're talking about, but in a head-to-head benchmark test you'll find results in the ballpark of 2-10x CPU and 0.5-2x GPU. As you point out, the power and cost are big differentiators. The other thing to consider is (as another commenter mentioned) that usually inference on CPU or GPU will require you to do some model quantization or compression, which can degrade model accuracy. Tensil can give you a way around that dilemma, so that you can have great performance without sacrificing accuracy.
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Thank you for the kind words! Just to clarify, the core technology here is free and open source, anyone can use it right now for free. We do have commercialization plans in addition - we may explore things like additional paid features for enterprise use or paid tiers of extra support.

Regarding LSTMs, yes. We're aiming to support all machine learning model architectures: do you have any particular models you're interested in that we should be prototyping with?

For CGRAs, we don't have any immediate plans to explicitly support them. What kind of use case do you have in mind? Generally, any platform that can implement a blob of generated RTL should be something we can work with quite easily.

tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
It depends on the model, yes. Here are some examples in the benchmarks section of our docs: https://www.tensil.ai/docs/reference/benchmarks/

We haven't specifically tested on any ICE40 FPGAs yet - if this is something that you'd really like to see, let me know! Taking a look at the lineup, the ICE40 LP8K and LP4K would be suitable for running a very small version of the Tensil accelerator. You'd want to run a small model in order to get reasonable performance.

Generally speaking, FPGAs with some kind of DSP (digital signal processing) capability will work best, since they can most efficiently implement the multiply-accumulate operations needed.

tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
You absolutely can use it in a data centre. You can even tape out an ASIC using these designs! Currently we've done most of our prototyping with edge FPGA platforms but if you want to try other platforms we'd love to help you get started. You can email me at tom@tensil.ai or use the contact methods on the website.
tdba··on Launch HN: Tensil (YC S19) – Open-Source ML Accelerators
Great question - TVM / OctoML are a great option if you have an off-the-shelf ML model and off-the-shelf hardware. Tensil is different in that you can actually customize the accelerator hardware itself, allowing you to get the best trade-off of performance / accuracy / power usage / cost given your particular ML workload. This is especially useful if you want to avoid degrading the accuracy of your models (e.g. through quantization) to achieve performance targets.
tdba··on Is Culture Stuck?
Tl;dr

> Culture is no longer made. It is simply curated from existing culture, refined, and regurgitated back at us. The algorithms cut off the possibility of new discovery.

tdba··on Procrastination is driven by our desire to avoid difficult emotions, says expert
Well said.
← PreviousPage 4 of 4