Senate passes bill to decrease grid digitization, move toward manual control
utilitydive.com
utilitydive.com
Alternatively, this could also be a misguided effort that holds back smart-grid tech in America by a decade. It's really hard to tell.
While I have no clue what the security or reliability looks like in industrial applications, I would hope they evaluate their solutions with reliability in mind.
It's barely better than iot stuff. And some "older" systems probably still need NT4 to run (I mean, if you're lucky)
Keep hoping... The number of industrial systems deployed that are connected to the internet with the default passwords unchanged or no passwords is nauseating.
https://threatpost.com/exposing-scada-systems-shodan-110910/...
How have they not been exploited/fixed already? Shodan's made it clear for years, of course anybody can use zmap/masscan.
Without having hacked a database of patient to device mappings, it's not like you can ask for ransom.
Also, you may already want someone in particular dead and are just looking for a vector. This seems almost untraceable.
I'm thankfully not this person, but at a basic level I could see some bored teen turning off a dam just for fun.
Honestly the only reason I don't is because it would be "wrong" to do, but the internet is full of people, it's certainly not a stretch that somebody would want to.
It does come with some vulnerabilities. I'm not sure what the best way forward on those is; I suspect this bill's effects, if any, will not result in a more resilient or less vulnerable grid system. Attackers will just find different leverage points to attack.
I live in the Finnish countryside, and since much of the power grid out here is still ancient, there are power outages rather often.
However, if the power goes out, there's always something that tries to power everything back on automatically after exactly 30, 60, 120 and 240 seconds. If something's still wrong, the power will simply flick on and then off within a few seconds. If an outage lasts more than four minutes, it usually takes much longer to get it running again (I assume it's because it then triggers an alert that will be handled by a human operator).
I would assume there's a lot of similar stuff going on in the US power grid, considering it's immense size what I suspect is a somewhat similar mix of new and very old equipment.
And yes, surge protectors are mandatory here.
Sounds like a recloser, which tries a number of times to restore power under the assumption that most faults are transient (for instance, a branch falling in the wires and then sliding to the ground). After a set number of retries, it gives up and stays open; at that point, a human has to locate the fault, clear it (for instance, poking the fallen branch with a stick until it falls to the ground), and then tell the recloser to close the connection again.
I read a statistic at some point (which I can no longer find the reference for) that power outages tend to be bimodal: either they last a short time (e.g. <5min) or they last a long time. There's hardly anything in between.
In the US, utilities report outages longer than five minutes:
Apply brakes while we work on the steering. This is good policy. “Internet Of Shit” is no joke.
We reported it as fast as we fucking could. It took about 6 hrs for them to get it off the public internet with no password.
From our assessment, it looks like that was a complete mistake. And those do indeed happen. But what if someone were to have found it and were malicious, and started randomly mashing buttons?
That's why our jobs as sysads are never done. We occasionally do boneheaded stuff. Users also do. And so do contractors. And I'll do what I can when I see something wrong. Professional courtesy.
If there's a reason for using an electronic system, by all means.
But for God's sake air-gap it, compartmentalize functionality with clearly defined interfaces, build simple fault-tolerant monitoring, and make final controls hand-on-metal.
When multi-million dollar warships are crashed because of poor UX, I have zero faith we're safely managing complexity.
My understanding of the incident is that the auto-detection functions failed and when human functions were not done per SOP due to exhaustion and senior command elsewhere the vessel came into jeopardy. Then when inexperienced and untrained crew members tried too late to evade. An apparently too complex UX was sending ship controls to more than one console. This culminated into one of the multi-million dollar warships crashing.
You may have missed the opportunity to experience this distinction first hand, but I can assure you with the utmost confidence, the meta does indeed change.
[0] https://en.wikipedia.org/wiki/USS_Fitzgerald_and_MV_ACX_Crys...
https://medium.muz.li/ui-confusion-almost-sunk-a-us-destroye...
My $30K car does a fair job of nudging the steering wheel if it thinks I'm running off the road, surely a $2B battleship could have some basic collision avoidance capability too.
To my knowledge the detection systems (for detecting ships ect) were either not calibrated/failed making a ship impossible to detect. There may have been a serious lack of training spotting sensor failure too. The ship was due for repairs and had been put off. Also, according to SOP there should have been someone making rounds and looking out and identifying ships. Because honestly how do you miss a huge cargo ship lit up like a Christmas tree
There were multiple critical failures leading to the crash. Bad training, poor maintenance, and lax procedural compliance.
this is simply not true. there are finite particles in discrete space so there's no infinity here.
Not because it isn't effectively beyond linear counting or coverage but because it describes how it grows - in different permutations. Especially with actual other actors involved and literal arms races. Sure your expensive plate armor may stop blades, arrows, and even current musket fire but would it stop cannonfire?
Why do you think flying is so safe? It's because of automation and redundant systems.
Do things go wrong? Yes. But imagine how many things would go wrong if you gave people a 15,000 line check list.
Humans should still pilot the ship, but if they are going to let a collision happen, then the ship should try to avoid it even without human intervention (unless they press the "I know what I'm doing and want to do it anyway" button after the ship sounds the alarm saying it's taking evasive action)
So the system doesn't have to handle all of the edge cases, just the "We're on a collision course with that ship ahead" case.
It's not that automation wouldn't work in the absolute majority of cases, and that we couldn't define the cases where automation will have to know it cannot make it.
It's not that humans would always be better than automation: humans make mundane mistakes and casual accidents happens, and humans make bad judgements also in extreme situations.
Looking at the catastrophes in aviation and shipping from the recent years the problem seems to lie where automation and human controls interact.
In the typical case the automation is not completely switched off while the human still tries to take control of the system and some override functionality still triggers the automation without the human realising that.
Maybe there's a sensor failure that causes the automation to misbehave. Then the human controller turns off the autopilot and begins to apply the proper fix but the electronic system, through its faulty sensors, sees it differently and concludes that the human is doing something way too dangerous and doesn't let him. Then a tug of war begins with the human controller where the automation is thinking that the situation is dire and the ship must be saved while the human realises that something is not working properly but in the heat of the moment doesn't have time to diagnose or narrow down to shutting off automation and just keeps applying more.
If only the automation knew it can't "see" properly and the human is already trying to fix things, or if only the human knew part of the automation is still active and secretly, with no indication, countering the proper corrections by the human controller.
Electricity, radiation, remote or sealed mechanical systems, deeply layered software.
In all of these, you need systems to communicate actual state to a human to effect even a manual resolution.
If a human is in the loop at all then it's because they may be expected to make a decision. And that decision is based on information. And that information needs to be presented unambiguously.
One of the smartest things I read on feedback here was to start with a balance-in, balance-out counting system.
It's unambiguous: if the number of things being processed and output is different than the number of things input, that tells me something.
And that's where automated systems get dangerous. It's incredibly risky if instead of informing and shrinking their response space, they expand it in an attempt to auto-correct the issue.
A quick google turned up this fairly lengthy page https://www.electricaltechnology.org/2018/03/electrical-prot... but the protective systems are covered in depth in basic power transmission tech texts. Also here https://en.wikipedia.org/wiki/Recloser
Why don't we have a proverbial two key system for certain activities like deploying to production or deleting user accounts, or backups?
We make do with things like pull requests, but everywhere we've ever had retros, I've run into situations where someone did something very stupid, somebody else rubber stamped it, and now we have a mess. Approvals on an action do not approximate the ritual sobriety that is embodied in two independent humans deciding to turn a key.
There are some people for whom these two actions can be considered equivalent, but we tend to bag on them for their meticulousness.
I think I'm wanting a tool where I type "deploy 1.0.12345 to production" and it sits there waiting until someone else types "deploy 1.0.12354 to production", and then it stops us because one of us transposed some digits.
It's actually a plugin for sudo (surprisingly yes, sudo has plugin capabilities[1]) and not PAM. I had originally developed it as a PAM module, but the sudo plugin API allows for the neat trick where the TTY is mirrored.
https://github.com/square/sudo_pair
In action: https://raw.githubusercontent.com/square/sudo_pair/master/de...
And to stouset: that's a seriously cool project!
It might, but policy could specify not allowing interactive mode for situations that dictate a two-man rule. Violation of said policy would necessitate a review and possible HR-related action depending on how serious the company wants to treat things.
If software were to become reliable, I'm sure those pioneers in reliable software can pave the way to reasonable use in mission critical, high value espionage targets.
Until then, legislation like this will pave the way back to what has worked for centuries.
You also need some connection, somewhere, to the “regular” internet. As just one example, consumption data is the basis for billing, which needs to connect to financial infrastructure.
And, finally, even air-gapped systems are vulnerable. See Stuxnet.
Say everyone has a registered nonrepudiatable cryotographic certificate - a decent idea in theory. However apply it to every accessible grid communication part including ones which may cycle in and out and could have secrets extracted. Suddenly it doesn't provide as much benefit if a hacked SmartMeter can get in anyway and there is deniability. Better than nothing but you could have likely gotten better bang for your buck with strict firewall rules or hiring a swarm of consultants.
Similarly vetting commands has two problems - one is that reaction time may be critical. Second the judgements themselves are complexity and may be a source for errors - especially when they are "triage" ones that cause a shutdown or damage to avoid greater damage
An air gap didn't help Iran against Stuxnet, or the US DoD against agent.btz:
* https://en.wikipedia.org/wiki/2008_cyberattack_on_United_Sta...
It wouldn't hurt, and it's probably worth doing anyway as a risk mitigation, but keeping it completely sanitary would be impossible.
The headline seems to be inferring more than the content is saying. The intent is not to rollback digitized systems, it's to decrease dependency on vulnerable systems. And the bill is only to commission a study on what systems could benefit from manual fallback, nothing is slated to be implemented yet.
Increasing renewable penetration into a power system leads to:
-> reduced inertia
-> increasing Rate of Change of Frequency (RoCoF)
-> reduced time to react
If you rely entirely on manual systems to react to frequency changes on a grid with high renewable energy penetration (e.g. the Irish grid), you'd really struggle to avoid daily blackouts without decommissioning a bunch of wind farms.
It is by using such automated control schemes that countries like Ireland are able to push the limit on wind power penetration.
If I was being truly cynical, I would think that a push towards mandating manual control (rather than mandating that it is available as a fallback) is actually a way to limit growth of the renewable energy sector.
but i’d think they’d be largely automated to reduce human error?
My point was that a grid comprised of a higher percentage of nuclear wouldn't need as advanced networked feedback, as sources wouldn't be turning on and off suddenly.
Indeed there would be local digital monitoring and control with nuclear, but that's a much less risky threat vector.
There is no valid reason to make nuclear power plant controls accessible for control by internet.
Nuclear is statistically the safest form of energy in terms of deaths/kWh and waste storage is massively overblown, not even accounting for how we could be recycling that waste in new plant designs if it weren't for the fubar politics around it.
This is somewhat misleading, because the major risks are things like birth defects and cancers which are really hard to attribute to any specific cause. Not to mention that it's stochastic; we probably won't have a disaster, but if we do it ruins a huge land area at tremendous cost.
Also, this doesn't work for every country.
For example you could build an array of flywheels that charges when the frequency is above 50Hz and discharges when the frequency is below 50Hz.
No network connection necessary, and if you want to you can build it entirely analog. If all you really want is inertia and reaction time it could even be as simple as dumb three phase motors with weights.
All rotary field machines (power plant generators, phase shifters, motors) work this way and contribute stability to the grid.
These basically add inertia (i.e. dampen frequency changes), but they do not regulate frequency. That's a fundamental difference.
But really inertia is all you need because all we are trying to do here is replacing the lost inertia from replacing spinning generators with solar. The recovered inertia gives human operators the time to make phone calls to increase energy production, spinning up a pumped-hydro plant or whatever is available.
I'm not sure if frequency regulation is currently considered an ancillary service in electricity markets right now though - from what I know about NY state, only 10- and 30-minute reserve are priced.
As someone responsible for the IT systems reliability of a grid operator I have to wonder if this really is a backhanded attempt to limit to the growth of renewables.
We run load frequency control on the grid every 4 seconds to keep the grid in balance. Hard to see how that'd work with manual controls on today's complex grid.
https://www.energy.gov/articles/department-energy-announces-...
Physical grid improvements may be as important as computational ones. More available and cheaper electricity is crucial to reducing the use of gas-powered building heat systems, which represent a large chunk of CO2 emissions:
The old phone networks were analog and vulnerable to cereal box whistles.
We are clearly doing something right, and automated systems replacing error-prone humans may well be a part of it. I don’t quite get HN’s infatuation with the mythical perfect pilot.
Stuxnet successfully attacked air gapped systems.
A system which "should" only allow read access is a system that will most likely have flaws which allow for far more, especially given the governments tendency to underspend on salaries for the people responsible for implementation.
I'm a fan of manual systems, but within reason though. The point is, you're never going to be 100% protected. People can still break into substations and compromise the grid rather easily.
they don’t really have write access, but they kinda do?
An Internet rando should not have the ability to disconnect your 700kva interconnect over the public internet. If you breach Tesla and can command enough vehicles to charge in a constrained area to overload local distributed infrastructure, failsafes should kick in and physically segregate loads from the local grid.
Disclaimer: This is only my opinion as an infosec practitioner having done some infosec/GRC consulting for investor owed and coop utilities for FERC compliance.
a) Would you send a lineman out to the recloser and have them manually configure it, spending probably hours doing so?
b) Or would you have some operator press a button in a piece of software to engage the recloser (thus, ultimately, writing whatever register in the device to perform the closing)?
Now multiply that scenario to the potentially hundreds/thousands of reclosers in the field.
Companies are going to choose b) every time.
If you want to interrogate recloser behavior remotely (real time and for historical logging), entirely reasonable. If you want to reclose manually or update trip and reclose thresholds, your commands or setting updates must be authenticated, logged, and cryptographically signed. If the ACR encounters lockout, you roll a truck (which you're probably going to do regardless, as lockout indicates a non-transient fault).
The only real vector one has into such a system (assuming you can't physically access it) is by attacking the 'inputs' to the SCADA network.
Attack vectors like stuxnet or petya require physical access to these systems (i.e. a usb key getting plugged into an operation terminal).
I think you’ve got this backwards. This is exactly what is done _in theory_. In practice many of these systems are not. When I was working as a consultant, I saw so many exposed SCADA systems.
It was a popular tech story a few years ago.
Mandate paper ballots next please.
Consider that the way we currently interact with power plants: "Someone needs to stay behind and manually make the boilers explode! I'll do it! Tell Laura I love her." pales in drama to "OK, let's see, Power / Input / Boiler / Max, confirm, confirm, confirm. OK, let's run!"
Stuxney penetrated a heavily guarded, monitored, secured with air gap facility. US power grid systems are child’s play in comparison
I don't see the problem with striking a balance with technology and putting in low-tech solutions for security reasons. So yeah I think many of those systems should implement similar measures.
This is quite an indictment of the security community.
Even if a bad actor did bring down the grid all it would do to concrete capabilities is GDP damage that they could easily get credit for and still leave a pissed off superpower able to make an example. Even actual rival status like say China were to do it would be shooting themselves in the foot such that something stupid and self destructive that if the PRC decided to simply start bombing Beijing it would do less damage.
Besides for decades a few tens or hundreds of people could go out, buy semi automatic rifles and start casually shooting substations and walking or driving away until someone arrests them. That would cripple power infastructure far worse than any hacker could as it would take longer to replace and be a distribuited issue to fix.
Given actual frequency even with lax security and pessimistic assumptions efficiency is a likely winner. Logistics and not superweapons win wars.
Plus the current track records of internet security for physical control systems is terrible.
Just look at how some companies interview and select people. They do so on the basis of cleverness, not carefulness and attention to detail.
You know Java? Cool. My kid is learning that in high school.
That's a direct reflection of how we develop software. New and shiny trumps reliable and boring any day of the week. That, coupled with an industry that does not even recognize the concepts of liability and defective product is a surefire recipe for a very large disaster at some point in the future. And after that it will get a lot easier to get budget for security, testing and all those other things that companies see as unnecessary cost.
There are also plenty of reasons to automate things. It allows for for faster reactions to problems and interesting new ways to reroute power. Both approaches have pros and cons, and there are good reasons for backing either approach.
Russia attacked Ukraine in a new way, and we responded by trying to become more well-defended against that new attack. Much as I agree that coding interviews are problematic, I fail to see how your point follows.
Digitizing the grid has enormous upside, to the tune of billions in savings and improved resiliency/response to weather related outages. It we could do it safely and securely it'd be a no-brainer.
We're just trading a devil we know for a (preventable) devil we don't.
BTW: digitized grids aren't even necessarily more vulnerable. In a complex system, the increased latency and miscommunication opportunities introduced by human operators are also a potential attack vector...
If you've ever watched your voltage, you'd noticed that it isn't a perfect 110 or 220. It is often higher or lower. When it is higher, there is a local surplus, when it is lower, there is a high load.
We could do this today. We might not have current pricing, but we do have load vs production information.
Or perhaps the voltage got too low, and an on-load tap changer in one of the transformers increased the output voltage. Voltage does not necessarily follow the load. AFAIK, the thing generators themselves use as the main feedback signal is not voltage, but frequency; but it's not a useful signal for consumers, given that generators are much stronger at keeping the frequency at its nominal value.
Load causes the voltage to drop (that's what's happening when a "brown out" is triggered). Some loads cause the current to lead or lag the voltage wave (Inductive vs Capacitive loads, most are Inductive, particularly with heavy duty equipment). But that isn't changing the frequency but rather the phase of the current. This is all tied up with a number referred to the "Power factor" (see https://en.wikipedia.org/wiki/Power_factor ). essentially, the farther shifted current is from voltage, the more work is done by the power plants essentially heating grid wires (rather than doing something useful)
So, power grids will do 2 things. First, they'll work to keep the current and voltage phase in sync. They do this by adding extra capacitors/inductors.
Second, they work to maintain the voltage of their tie in to to the grid.
Generally speaking, the type of power plant matters as well. Base load plants will simply dump onto the grid at a constant rate (without really caring about what the voltage is) while peaker and load following plants will attempt to vary output relative to their voltage to try and keep the grid voltages stable.
You are correct, the voltage variance can be misleading at the customer level if the transformer is actively adjusting it's voltage ratio. I didn't consider that.
And we've done it before (safe, anyway). Is our electric grid as critical as space shuttle software?
NASA’s Mars Climate Orbiter; crashed or is now inoperable and orbiting the sun due to a bad numerical conversion.
ESA’s Ariane 5 Flight 501; manual self destruct triggered after a 64bit number being truncated into 16bits caused faults to be thrown.
https://raygun.com/blog/costly-software-errors-history/#
I don’t think we should go back to manual operation (as the default, but should be overridable). Instead we should be using stricter compilers, and better unit, integration tests, and fuzz testing to test as many edge conditions as possible.
It does take both kinds of kinds... creative thinking for hard problems, but attention to detail towards the path of correctness. I completely agree the "cleverness" is massively over-emphasized during the interviews for most companies