112 karma · joined July 1, 2015
Automox has seen double digit growth year over year, and this year is no exception. We recently raised $110 million Series C led by Insight Partners and appointed Dmitri Alperovitch, Crowdstrike co-founder as the Chair of our Board. We are modernizing the IT Ops market with our cloud native approach to automating and streamlining IT workflows.
With over 100 positions to hire for this year, we have something for everyone. From building data pipelines, scaling our infrastructure in AWS, developing distributed backend services, to simplifying a complex UI.
Our tech stack = AWS, K8s, Kafka, RabbitMQ, Java, Go, Vue.js
Hiring:
* Staff, Senior, Mid Level Engineers across the stack * Senior Data Engineers * SDETs (Python) * Windows or Linux or Mac Systems Engineers * UX - Researcher and Architect * Technical Product Managers * Senior Site Reliability Engineers
If you are interested, apply here: https://jobs.lever.co/automox?lever-via=qv0yUNv1SS
* Use a monorepo to "namespace" different projects/teams/whatever. Each namespace has its own build.sbt for Scala jobs and Conda/Pip requirements file for PySpark. This gives you package isolation so that different projects can bump requirements at their own pace. This is crucial in larger organizations where you might have more siloed development or more legacy applications.
* Build each project in the monorepo into a separate Docker image and tag it accordingly with some combination of the branch and namespace.
* Deploy applications onto Kubernetes by invoking the SparkOperator (https://github.com/GoogleCloudPlatform/spark-on-k8s-operator), This abstracts away a lot of the hassle of driver/executor configuration and gives you nice out-of-the-box functionality for scraping Spark metrics.
* For local development, use some type of CLI or Makefile to build/run the image locally. This is where the implementation diverges somewhat from using SparkOpelrator (unless you want to tell your employees that everyone needs to run Kubernetes on their local machine, which we thought would create too much friction).
* For orchestration, write a custom operator for Airflow that submits a SparkOperator resource to the Kubernetes cluster of your choosing. The operator should supervise the application state, since the SparkOperator doesn’t quite do that well enough for you. This is something I wish we had the opportunity to open source.
* Where it gets tricky is building Spark applications locally and running remotely, Say you built a job locally and tested it on a small subset of your data. Now you want to see what happens when you run across a full dataset, requiring more than 16gb of memory (or whatever the developer has on their laptop). You need some way to build your image locally but schedule it remotely. This could be done via the same CLI or Makefile, but you end up with a lot of images and it gets pretty costly. I’m sure we would have figured it out eventually if we didn’t all get laid off last month :P
* BONUS: Use Iceberg or Delta (https://iceberg.apache.org/) (https://delta.io/). These are storage formats that work with distributed file storage like HDFS or S3 to partition and query data using the Spark DataFrame API. You get time travel, schema evolution and a bunch of other sweet features out of the box. They are an evolution of Hadoop-era partitioned file formats and are an absolute must for organizations dealing with lots of data & ML infrastructure.
This post took up more time than I had wanted, but it actually feels good to write down before I forget. I hope it is useful for someone building Spark infrastructure. I'm sure others have a completely different approach, which I'd be curious to hear! As someone whose full time job was basically just to orchestrate Spark application development, I can say for certain products like this are needed in order for the ecosystem to thrive, and I would probably have given you my business had the circumstances been correct. Good luck to you and your team.
https://medium.com/fluidity/keyspace-end-to-end-encryption-u...
Or, we could try to shift agriculture back to a local scale, use open source hardware/software, and community-owned infrastructure to build more sustainable, polyculture food systems.
In particular, I am excited about the rooftop farming work being done in the Brooklyn Navy Yard here in NYC(1). Our Public Advocate has even discussed building-code mandated "green roof" legislation(2). CNC/Robotics & IoT are the key to unlocking urban micro-agriculture that can begin to offset some of our dependency on dirty food, and I applaud those(3) who are working on these very important problems.
(1) https://www.brooklyngrangefarm.com/navyyard
(2) https://www.bkreader.com/2018/07/18/brooklyn-councilmember-i...
Our schema involved taking physical assets/personnel and representing them as different labels: machine, factory, production line, user, usergroup, etc. We then drew complex relationships between different user/groups in the organization and the assets they were responsible for.
At first, we used a relational database, but it soon became difficult to go more granular than simply: user belongs to usergroup, usergroup belongs to client, client has factories, factories have lines, lines have machines.
As many have pointed out here, it's not that you can't do this with non-graph databases, it just requires a more complex query layer. Neo4j allowed us to represent complex business relationships as natural language, and that really helped us as the business scaled.
https://www.nytimes.com/2018/07/27/us/politics/russian-hacke...
I'm willing to bet money this was a cyber-terrorist attack. Unfortunately we'll never know. If a link were established, it would be the subject of a gag order on grounds of national security. But more likely, the true root cause will never be found because the authorities didn't do a deep enough forensic analysis. It's too easy to blame something this on mechanical failure, especially in America's aging infrastructure. They won't even think to look at the PLCs and control systems that control the gas pumps :/
UnsupportedAvailabilityZoneException: Cannot create cluster because us-east-1b, the targeted availability zone, does not currently have sufficient capacity to support the cluster. Retry and choose from these availability zones: us-east-1a, us-east-1c, us-east-1d
But yeah, sure, keep telling yourself the AWS doesn't have a problem with power and compute capacity. Or maybe it's just poor product design?
https://docs.aws.amazon.com/eks/latest/userguide/troubleshoo...
They should've called it Generally Available(ish)
We're a small but rapidly growing team focused on building products that allow manufacturers to improve their production processes using data. We’re working across a range of cutting edge disciplines including industrial Internet-of-Things, big data, and machine learning.
We have openings across the board:
- Frontend: help build the next iteration of our manufacturing analytics platform, a first of its kind suite of applications for analyzing real-time data, optimizing production processes, and modeling the factory of the 21st century.
- Backend: build highly available APIs in Python / Go that efficiently and reliably capture machine and human data.
- DevOps Engineer: tackle interesting problems with infrastructure in a hybrid cloud & IoT environment, such as quorum-based distributed systems and cloud/edge application deployment strategies.
- Data Scientist: build statistical and machine learning models that improve efficiency of manufacturing using the telemetry collected from machines in the field.
- Forward Deployed: integrate with different production machines, allowing for seamless transmission of data to our platform.
- Customer Success Manager: ensure that our clients are using the product to achieve the best possible outcomes for their business. This person is ideally an operations/logistics/industrial consultant, engineer, lean expert, or similar with a proven track record for demonstrating ROI.
Reach out directly: mykola [at] oden [dot] io
From Wikipedia: "Generally substations are unattended, relying on SCADA for remote supervision and control." These SCADA systems may or may not be connected to Internet, which could allow an attacker to remotely access and modify the code that controls transformers and other electrical equipment.
In another comment, someone mentioned the power company was, in this case, pumping C02 into the substation in order to contain smoldering electrical insulation. This means that, most likely, the copper conduit heat up beyond defined tolerances. This could be due to more current being carried than those conduits are rated for. Normally, the SCADA system would be responsible for keeping these currents within tolerances. What I am saying is that they could deliver a payload to the PLCs via the SCADA that could trick the transformers into taking more load than they could handle.
We are an IoT startup creating a hardware / software platform for Industry 4.0 [1] factories. We collect data from industrial machinery and analyze, aggregate and display it so that manufacturers can make more product with less material. There's a lot of exciting things happening at the company and now is a great time to get into a small (8-person) team working working on a lofty mission that will revolutionize an underserved industry.
We're looking for a data engineer with experience in building realtime and batch processing data pipelines. We ingest tens (soon to be 100s) of millions of data points daily and do complex aggregations and calculations that help our customers to hone their manufacturing processes. If you have experience with lambda architecture, timeseries / graph dbs and cutting edge data engineering technologies, we need your help ASAP.
Feel free to reach out to me directly: mykola@oden.io
Point is, we're a hybrid shop that does all first pass services in Python 2.7 and then move them to Go when they become suffiently trafficked and/or critical.
We are an industrial IoT company that allows manufacturers to optimize processes and produce more output with less input by improving efficiency and reducing waste products. Our goal is to create smart factories using cutting edge technologies. We are currently funded, w/ a small # of employees. Now is a great time to get in ;) Stack: Python, React, ConcourseCI, Cassandra, KairosDB, MongoDB, Go (nothing is set in stone, we value engineers that take a scientific approach to evaluating all possible solutions — help us decide!)
* Hardware & Network Engineer: https://odentech.recruiterbox.com/jobs/fk06pfp/ We need engineers, preferably w/ experience in IoT, to help us build out our hardware and network strategy. This includes writing software for embedded devices, experimenting with different network connectivity solutions, and optimizing device firmware for reliability and security.
* Frontend Engineer: https://odentech.recruiterbox.com/jobs/fk06s2b/ Our end-user product is a dashboard that allows factory workers to grok massive amounts of timeseries data. Experience in analytics or visualization solutions is preferred.
* Data Engineer: https://odentech.recruiterbox.com/jobs/fk06s2k/ We currently ingest 8.5M datapoints per day and expect that number to increase 100x by the end of the year. We are looking for a skilled big data engineer to help us ingest and process this data.
* Backend / Realtime Stream Processing Engineer: https://odentech.recruiterbox.com/jobs/fk0hdsz/ Imagine two machines reporting datapoints at different intervals, that are components of a complex aggregated metric. We need to be able to perform aggregations on datapoints as they arrive, in as close to realtime as possible. Experience in realtime stream processing libraries and out-of-order event processing is a plus.
Feel free to apply on Recruiter Box (make sure to mention HN), or reach out directly: mykola@oden.io
We are an industrial IoT company that allows manufacturers to optimize processes and produce more output with less input by improving efficiency and reducing waste products. Our goal is to create smart factories using cutting edge technologies. We are currently funded w/ a small # of employees. Now is a great time to get in ;)
Stack: Python, React, ConcourseCI, Cassandra, KairosDB, MongoDB, Go (nothing is set in stone, we value engineers that take a scientific approach to evaluating all possible solutions — help us decide!)
* Hardware & Network Engineer: https://odentech.recruiterbox.com/jobs/fk06pfp/ We need engineers, preferably w/ experience in IoT, to help us build out our hardware and network strategy. This includes writing software for embedded devices, experimenting with different network connectivity solutions, and optimizing device firmware for reliability and security.
* Frontend Engineer: https://odentech.recruiterbox.com/jobs/fk06s2b/ Our end-user product is a dashboard that allows factory workers to grok massive amounts of timeseries data. Experience in analytics or visualization solutions is preferred.
* Data Engineer: https://odentech.recruiterbox.com/jobs/fk06s2k/ We currently ingest 8.5M datapoints per day and expect that number to increase 100x by the end of the year. We are looking for a skilled big data engineer to help us ingest and process this data.
* Backend / Realtime Stream Processing Engineer: https://odentech.recruiterbox.com/jobs/fk0hdsz/ Imagine two machines reporting datapoints at different intervals, that are components of a complex aggregated metric. We need to be able to perform aggregations on datapoints as they arrive, in as close to realtime as possible. Experience in realtime stream processing libraries and out-of-order event processing is a plus.
Feel free to apply on Recruiter Box (make sure to mention HN), or reach out directly: mykola@oden.io
Like many other users here, we were disappointed about paid clustering, but when the original press release said $400, we were willing to wait it out and see. However, we ultimately decided to go a different direction after seeing they wanted $20k+ to run clustering on a 256GB node. We ingest 10s of millions of data points per day from IoT sensors, and expect our data size to far exceed that capacity.
That said, we plan to run our own Cassandra cluster w/ KairosDB (http://kairosdb.github.io/) acting as a read / write abstraction layer. It'll cost us about $11k to run the cluster for the year, with 3 nodes @300GB/ea., leveraging Cassandra's (free) and open source clustering, HA, and replication technology.
We use a combination of proprietary and open-source hardware+software to gather data from industrial machinery and push it to our platform wirelessly so that manufacturers can analyze and optimize their production. I was the first external hire for the company and actually found out about it from a "Who is hiring?" thread! We're working on some really interesting problems that will require traditional and out-of-the-box thinking and need a super talented DevOps lead to get our infrastructure to where we need it. Feel be free to message me with any questions (mykola [at] oden.io).