Full disclosure: I'm an employee
16,818 karma · joined February 17, 2009
http://druwynings.com
Feel free to email me about anything: dru@druwynings.com
Head of Marketing @ Sensible.so. Previously Head of Marketing and Business Development at Diffbot (http://www.diffbot.com/), BizDev @ Heyzap (YC W09). Self-taught programmer & designer.
@druwynings
Full disclosure: I'm an employee
Sensible is an API-first document processing platform that streamlines data extraction for developers and product teams.
We're looking for an experienced Product Marketer to join our growing team. As a product marketer, you will lead core product marketing initiatives, build awareness with our target market, and generate inbound leads for our sales team.
This role requires a mix of creative and quantitative thinking. You'll develop a product messaging and positioning framework that resonates with developers and presents our products in ways that increase Sensible’s ACV.
Ultimately, you will drive awareness, product adoption and revenue growth. This position reports to the Head of Marketing.
Sensible is a remote-first company and this role is open to anyone located in North or South America. We provide competitive compensation and meaningful equity.
More info: https://sensiblehq.notion.site/Technical-Product-Marketing-M...
-- If this sounds interesting, please reach out to dru@sensible.so
Sensible connects software to the messy, unstructured data that businesses encounter every day. Our goal is to bring about a world where computers do the work that computers are best at (processing large volumes of data) so that humans can do what humans are best at (critical thinking and empathy).
The first piece of this problem that we're tackling is document parsing. Even as software is eating the world, so many business workflows still rely on two entities sending PDFs to each other. In a single afternoon with Sensible, developers can ship a production-ready API endpoint that turns documents into useful data.
Sensible was founded in 2020 and is backed by some of the top investors in Silicon Valley. https://www.sensible.so/about
--
As a senior infrastructure engineer, you’ll work closely with our head of engineering and the engineering team to improve and expand our AWS-based infrastructure, working with technologies like Lambda, DynamoDB, S3, and IAM, to support our production APIs and web app. Outside of AWS we integrate with Google Cloud, Microsoft Azure, and OpenAI. We have a strong testing culture to support our overall reliability, as well as SLAs for our enterprise customers.
--
You might be a fit if...
- You have 5+ years of experience building infrastructure and APIs in AWS.
- You’re product-minded and customer-focused.
- You have experience working in organizations compliant with SOC 2 and HIPAA.
- You are an excellent verbal and written communicator, able to build relationships with different kinds of people across different levels of the organization.
--
Interested? Reach out to the founder josh@sensible.so and mention HN.
--
- Senior Infrastructure Engineer: https://www.notion.so/sensiblehq/Senior-Infrastructure-Engin...
- Customer Success Engineer: https://www.notion.so/sensiblehq/Customer-Success-Engineer-8...
- Product Manager: https://www.notion.so/sensiblehq/Product-Manager-a0dc92f1a9d...
- Account Executive: https://www.notion.so/sensiblehq/Account-Executive-14205252e...
Answered on the parent, but it's somewhat similar.
Essentially that's what Diffbot (https://www.diffbot.com/) does, except we don't the render pages as an image nor do OCR.
Diffbot renders the page in a headless browser, and uses computer vision to automatically identify the key page attributes and extract normalized data for specific page types (Articles, Products, Discussions, Profiles, Images, and Videos).
This approach enables us to work in any language and on sites that we've never come across before automatically with better than human level accuracy.
If anyone has any questions or wants to try it out, feel free to email me directly at dru@diffbot.com
Major differences I can see (OP feel free to correct if I'm wrong):
Link.fish
* doesn't provide a web crawler
* relies heavily on microdata, schema.org, RDFa, etc
* relies on manual parsers for sites that don't have microdata embedded
* doesn't full-render pages by default (Diffbot renders every page, so it can use computer vision to automatically extract the data)
* doesn't support proxies
* doesn't support entity tagging
Probably plenty more, but that's what jumps out to me at first blush.
--
Since I see other people have mentioned price as a concern, we're always willing to help out bootstrapped startups. Just shoot me an email: dru@diffbot.com
A couple notes:
"the post-Duolingo learn language in context app"
This was really hard for me to understand. When I skimmed this, my first assumption was that this was a direct replacement for Duolingo, when it's actually complementary.It would probably be better to split up the long list of languages on to 2 different pages. 1. Language you want to learn, and 2. From which language. An alternative would be to have a separate section for each language to learn on the page with space in between.
Demo video: https://www.youtube.com/watch?v=g0O8UNM0B7Y
We're an AI startup that applies machine learning, computer vision, and NLP techniques to the problem of understanding webpages. Our APIs convert billions of webpages automatically into structured data for the likes of DuckDuckGo, Salesforce, Hubspot, Amazon, Bing, eBay, Adobe and others.
We recently announced our profitability(!!) and raised a $10M Series A by Tencent Ventures and Felicis Ventures.
Looking for ML/CV/NLP specialists, data fusion / knowledge graph, and/or web-scale crawling experts with a track record of building intelligent systems that perform at human-level accuracy rates.
AI Researcher: https://careers.jobscore.com/careers/diffbot/jobs/ai-researc...
Data Operations Product Manager: https://careers.jobscore.com/careers/diffbot/jobs/data-opera...
Search Engineer: https://careers.jobscore.com/careers/diffbot/jobs/search-eng...
Technical Account Exec: https://careers.jobscore.com/careers/diffbot/jobs/technical-...
If you have any questions, feel free to email me directly: dru@diffbot.com
If you'd like me to set you up with a demo, feel free to email me: dru@diffbot.com
All of our automatic APIs aren't meant to be used on directories or listings pages (for that, we provide a crawler which gathers those links and then feeds them into our automatic APIs). The discussion API is intended to be used on discussion pages, for instance: http://www.diffbot.com/testdrive/?api=discussion&url=https%3...
If you wanted just a list of submissions and score, you could use our custom API toolkit(similar to Kimono): http://www.diffbot.com/products/custom/
Sorry to hear about the issues you ran into when trying out Diffbot! If you have some examples, I'd like to take a look into whether we can improve.
In cases where our automatic extraction using computer vision isn't 100% accurate, we offer a visual interface for actually overriding our default extraction. This input is then used as additional training data for the ML models.
Mozenda et al. are great if you only need data from a couple of sites and you don't mind spending time manually specifying and maintaining CSS selectors for each website and page layout you need data from.
Our crawling, and proxy support, is fairly robust thanks to our hiring the creator of Gigablast [https://gigaom.com/2013/09/10/diffbot-brings-big-time-search...].
If you'd like to give Diffbot another go or you have some examples where the extraction could be improved, please let me know!
We're an AI startup that applies deep learning, computer vision, and natural language processing techniques to the problem of understanding webpages. Our APIs convert billions of webpages automatically into structured data for the likes of DuckDuckGo, Bing, Digg, Instapaper, eBay, Adobe and others.
We use very few 3rd party frameworks and strive to develop our own performant machine learning techniques.
Looking for published ML/CV/NLP specialists, data fusion / knowledge graph, and/or web-scale crawling experts with a track record of building intelligent systems that perform at human-level accuracy rates. If interested, send us a note at jobs@diffbot.com (due to limited HR bandwidth, naked resumes will be discarded).
If you have any questions, feel free to email me directly: dru@diffbot.com