"BGP at home": getting a DIA circuit installed at home
aaka.sh
aaka.sh
I currently have a BGP session with one of my ISPs on the 15 euro plan (1 Gbps), and another local ISP is giving me 10 Gbps for free because I'm a test user. Both of these ISPs support BGP, and it's amazing that I can have BGP sessions as a home user. However, it can be challenging to speak with regular sales reps who may not be familiar with BGP. It takes persistence and some effort to get to the network administrator.
It's frustrating that some ISPs make it artificially difficult for home users to obtain a BGP session, even when they're not asking for SLAs or dedicated connections. Technically, you can even run a BGP session over dial-up. It seems like these restrictions are in place to squeeze more money from customers.
It's no secret that we've run out of IPv4 addresses, and I happen to have a /24 subnet just for myself. Even with all my usage, I barely use half of it. There's no rule that says ISPs should only accept /24 or shorter prefixes, and this is contributing to the depletion of IPv4 addresses. I would happily announce a /25 and return the other /25 to the world, but I must have a /24 to do it. This practice of requiring /24 subnets may have made sense in the past to preserve router memory, but with today's technology, it's nonsensical and only serves to artificially deplete available IPv4 addresses.
Don't worry, it can be frustrating for businesses too !
Cogent famously nickle-and-dime their customers and consider BGP to be a chargeable extra, even on their IP Transit product which is somewhere where you would expect BGP to be a given.
Of course whether you should do business with Cogent is another matter, especially given their spammy cold-calling tactics, seemingly once you're on their list you can never get off it.
This is done via ASNs (Autonomous Systems Numbers), like 32150 (my old ASN). Each hop along the way gets their ASN tacked on. It's like a flood-fill to all the BGP routers in the world.
When you "traceroute", you see IP hops along the way, but the ISPs routers choose links based on this list of autonomous paths (among other things). Basically, a AS-level traceroute is built into the BGP routing information. At it's most basic, a path is selected based on the shortest AS path (traceroute, but at the organization level rather than IP level); the fewest number of networks the traffic has to traverse.
If a link goes down (either because of a network failing, or because of administrative reasons), traffic will eventually switch over to another way of reaching the destination. You can also do things like pad the AS path with multiple copies of your AS to cause traffic to switch to another link more frequently, as a kind of primitive load balancing.
There is really need to do this is many cases though. If all you are looking for is connectivity redundancy you can just take the default routes from two ISPs and configure floating static routes on your router and not worry about BGP at all.
>"This is done via ASNs (Autonomous Systems Numbers), like 32150 (my old ASN). Each hop along the way gets their ASN tacked on. It's like a flood-fill to all the BGP routers in the world."
Each hop does not get its own ASN tacked on. This only happens when crossing an AS boundary i.e at eBGP speaker. There are generally many hops inside an AS.
That only gives you outgoing redundancy. You can talk to the world with one of the links down, but the world can't necessarily talk to your IPs with the link down.
BGP lets you advertise IPs to the world and what links they're reachable on. You can announce the same IPs on multiple links to multiple ISPs. This is what makes the internet "route around damage".
Correct if we are talking about IP hops (router 1 to router 2 to router 3), the AS does get tacked on (possibly multiple times) for each AS hop (Level-3 to Cogent to Zayo). Thanks for clarifying that.
The TL;DR:
Yes, BGP allows you to advertise addresses.
Ergo, it is allows you to build resilience, i.e. you can connect to multiple independent ISPs and advertise the same addresses. Hence traffic will always be able to reach you even if one ISP goes FUBAR.
By the same token, it makes pretty much zero sense to have BGP on a home connection (or, in technical speak, a "single-homed connection"). Since really BGP is a zero-sum game on single-homed connections (no real benefit, much added complexity).
The primary benefit if fully-realized and is not about "is multiple paths via different carriers."
An AS is an administrative entity. Route prefixes belong to an AS. All routes with an AS have a common routing policy that is managed by a single entity. I can look up a prefix to find it's AS and then I can look up that AS's routing policies. Here is an example for Tier 1 ISP Spring:
https://www.sprint.net/policies/bgp
I could have single ISP that I buy transit from and I might have multiple links with them. If I want to understand how to take advantage of that for example with traffic engineering then I would look the policies for ISPs AS. I would then use those BGP community attributes on my BGP links to them in order to accomplish my own routing goals.
You should be surprised about how much of the internet still runs on gear from late 1990s
I also have a /24 for myself. I've had it since the mid-90's, from the InterNIC days, predating ARIN. I've gone years without using any of the addresses. Currently it's routed to my home lab. I'm doing BGP through a couple VPSes, then tunneling it back.
Basically, you get an ASN, find a BGP-friendly VPS or colo, announce your /24, and tunnel it back to your home.
>"Additionally, running one more BGP session makes their network appear larger, so they benefit as well."
This doesn't make any sense. The number of "BPG sessions" is not a metric any ISP thinks about. A BPG session is just TCP connection. The number of prefixes an ISP is advertising might be something they care about however but for smaller ISPs that doesn't really matter as they are just customers of the Tier 1 ISPs and that just requires paying the Tier 1 ISP. It has nothing to do with the number of sessions or prefixes. I could have a single /22 prefix for my company and the Tier 1 ISP will still sell me the same service.
>"There's no rule that says ISPs should only accept /24 or shorter prefixes, and this is contributing to the depletion of IPv4 addresses. I would happily announce a /25 and return the other /25 to the world, but I must have a /24 to do it."
Well yes there is. Many ISPs will generally not accept routes smaller than a /24. They configure filter-lists that prevent those routes from being accepted. The reason for this is to reduce the size of the global routing table. Each prefix has to be stored in TCAM(memory) on a router. There have been many incidents where older routers couldn't handle the number of routes in the global routing table and caused them outages. See:
https://arstechnica.com/information-technology/2014/08/inter... and https://www.inap.com/blog/growing-pains-internet-global-rout...
Further if you have a /24 you can certainly advertise 2 /25s to the same upstream ISP if you have more than one transit connection with them. In fact this is exactly how do traffic engineering.
>"It's frustrating that some ISPs make it artificially difficult for home users to obtain a BGP session, even when they're not asking for SLAs or dedicated connections."
This because home users absolutely don't need to run BGP! You just take a default route from your ISP and you're done with it. There is not point in running BGP if you are single-homed! You can't do any traffic engineering on it or influence the way upstream ISP route traffic to you. From the hobbyist point view it would be extremely boring. What can you actually do with it? View the global routing table? You can do this without ever running BGP by using a public looking glass server[1]. There is no incentive for ISP whose business is selling internet connectivity to cultivate a hobbyist community.
My home router is a MikroTik CRS309—it’s about $250, fanless, has 8 SFP+ ports and advertises my home network blocks over BGP to the two routers (for HA).
The setup works great. The best part is how it’s small and fanless and fits inside the very small wall box the fiber terminates in.
How do you fit two full BGP tables into 512MB of RAM? I've looked into MikroTik boxes before and maybe they're doing something I'm not understanding. On the routers I have manage, two IPv4 feeds take up about 1.1GB and three IPv6 adds another 450MB.
Except you do lose out on best path routing / any other outbound TE, and you’re now restricted to rudimentary load balancing methodologies / manual prefix-specific hackery.
For eyeball networks and other stub networks this is mostly fine.
For a home network I'm guessing you are pretty unlikely to do the multi-homing from the house, more likely you will either just have a single upstream or if you are connecting it into your own collocated infrastructure you will do iBGP and let your actual edge BGP routers handle the multi-homed upstreams and sync the full route table etc.
I was tempted to do this once before when I was running my own hosting company but it was prohibitive cost wise to get the circuit I wanted. :(
The parent I was responding to was explicitly claiming no downside to being default-only while multihomed.
Noction is a wonderful example of this. It has sane defaults, but insane customers, who think that it’s a good idea for them to originate more specifics for other outside networks, because they’ll never leak them outside of their own AS (narrator: of course, the prefixes leaked).
I think the most notable example of this recently was Verizon (the insane customer) using Noction (the SD-WAN technology) and doing exactly that, causing mass traffic disruption as they announced more specifics for other peoples prefixes, drawing all that traffic to their own network instead.
After 3-4 hops you're probably hitting a Tier 1 network, after which point you can basically think of the Internet like cloud icon you see in many diagrams, because your route choices are no longer really determining reachability, rather the choices of other people/companies are:
* https://en.wikipedia.org/wiki/Tier_1_network
If you're talking about reachability of a network on another continent or the other side of the planet, your local decisions aren't going to much to determine the path.
The main thing to have locally for routing decisions is the ASNs/networks of the other customers of your ISPs: if Service A is also a customer of ISP #1, you want to send traffic for them through that service instead of ISP #2.
The other nice thing to have is the reachability to the closest IXP, as quite often many CDNs have connections to those.
Beyond knowing IXP reachability and other-customers reachability, I don't think there are many other advantages for a smaller entity on the Internet, so a full Internet-wide BGP is not needed.
If they offer the option, instead of a "default-only" feed from each ISP, you may wish to see if they have "default-plus-our-customers" feed: if Service A is also a customer of ISP #1, then why bother sending the packets to ISP #2 in the first place?
* https://support.allstream.com/knowledge-base/bgp-request-inf...
In general, at some point you'll hit a Tier 1 network, after which it won't matter, but until that point getting the connectivity to other customers of the ISPs could be useful. The other reachability destination to pay attention to would be of IXPs, where CDNs often connect to.
bird> show memory
BIRD memory usage
Routing tables: 262 MB
Route attributes: 120 MB
ROA tables: 192 B
Protocols: 171 kB
Total: 382 MB
bird> show route count table r1
906312 of 906312 routes for 906312 networks
bird> show route count table r2
903532 of 903532 routes for 903532 networks
bird>
routes in kernel will take less as you're only getting best one exported to kernel and no route attributes to holdI've always considered Verizon's (and prior to that, WorldCom and/or MCI) 100% SLA to be pure marketing BS.
Anyone who's ever worked with telecoms knows it's simply unrealistic. Fibres will get cut, carrier's routers will need software updates etc. etc. etc.
I reckon the reason you pay a Verizon-tax is so they have spare cash floating around to pay you the inevitable SLA penalties.
Personally, I prefer dealing with carriers who have more real-world SLAs.
The whole thing has a lot of hack value, it's cool and also worth something to be in control of your own networking, even if it's more expensive. Like buying apple: overpriced for the specs you're getting, but you're assured it'll be good quality (that's the idea anyway) and it looks cool (to most). Except... apparently it's still got issues, regularly? Now I'm really wondering what the point was, at least with hindsight
There are real value in doing so in addition to the hack value, by doing so you can steer the Internet traffic to whatever IP blocks you owned to your house, dynamically. For example OP mentioned "add some resilience" as motivation, i.e. anytime their services running in the DC failed he can reroute the traffic to ... their house.
Lower bandwidth dedicated circuits can often "feel" faster than a higher bandwidth consumer connection.
Looking at past outages, I feel like given a certain budget, you would get much better uptime with two or more diverse (i.e. not "two fibers going exactly the same path") residential providers than with one commercial one, and for anything but the most critical projects, that's the way to go IMO.
Most importantly though, anything really really important that is so important to use custom fibers should be designed so it also works over the public Internet. Unless degraded service is worse than no service, it doesn't matter that the Internet is theoretically not reliable enough/doesn't provide enough guarantees - at least keep it as a backup so that when a backhoe takes out your custom fiber, you can switch over to the VPN instead of shutting down air traffic at one of Europe's largest airports.
Seems like another major reason for the decision here were SLO/SLAs. Those seem rather meaningless. The residential ISP will not try to max out their 95% SLO - an ISP that's down 1.5 days every month or leaves you offline for 18 days wouldn't be very popular. The commercial SLOs, on the other hand, sound great but they're not magic. If the issue can't be fixed in that time, they will be violated, and you'll typically get a credit for a few months of service - money that you wouldn't have spent in the first place if you just went with residential.
In the end, he still has countless single points of failure, e.g. the above-ground fiber line waiting for a tree, or a non-redundant power supply that he can't easily replace himself. Sure, he has an SLO, but with a dual uplink residential (which is cheaper), a CPE ("modem"/router) failure wouldn't be an outage in a first place.
The article says it's right off a major street. So is it urban? Suburban? Single-family home? Inquiring minds want to know.
Open-air racks are typically less than 2 feet across and 3-4 feet deep (example [1]), hardly a significant stretch unless you live in a very small space.
1: https://www.amazon.com/StarTech-com-Open-Frame-Server-Rack/d...
I went the route of “find a dirt cheap gigabit colo” ($150/m for quarter rack with gigabit and bgp peering) and just tunnel the IP space back home. Upside is you don’t have to pay to have the circuit put to where you currently live and you don’t have to worry about space/noise. Downside is if you want to use it locally at home you get extra latency due to the tunnel.
I guess it depends on how your tunnel is set up, but shouldn't you be able to add a static route to your rack on the "home side"?
I thought you meant "use the services hosted on your rack at home" (thus my suggestion), not "use the IP addresses you own at home".
It took me way too long to realize this option - many builds using miniture casings with cooling issues later, when buying large 2U cases with 120mm fan walls is the ticket.
> organizations design systems that mirror their own communication structure.
An org for the physical layer has a box, then the org offering the connectivity layer plugs their box into the first box. Like OSI layers stacking up.
Your assertion also seems contested by the words in this post:
> Dispatch 5: NID #2 install. This is where I began to see the distinction between Verizon Telecom (VzT) and Verizon Business (VzB). While the layer 1 infrastructure had been installed and tested, there still remained the IP layer, which required another dispatch for another technician to install a dedicated NID for Verizon Business -- a Ciena 3903, with 1310nm single mode optics.
It does mention a different fiber connection, but to me, this is describing each layer of the stack having both it's own organizational unit & it's own physical box that it manages.
Verizon is like a state or federal government agency. Massive complex bureaucracy that nobody really understands. As one of the ultimate legacy businesses they have all sorts of crazy settlements, agreements, tariffs, etc and segment their businesses around them.
Two different companies, one provides the last mile, one provides the customer service layered on top.
It adds a lot of flexibility, for instance in this case the person was on-net for Verizon, but what if they were not? Then it wouldn't change at all, just the first NID would be a different provider than Verizon.
The first device terminates the Fiber. It has an SPF transceiver in at for the wavelength the customer is getting the service on which is 1310nm. This is pure layer 1. A technician could/would use this "NID" to get a light reading during provisioning and/or troubleshooting. This is the end of the "local loop."
The second device is the one that provides Carrier Ethernet. Carrier Ethernet is like regular Ethernet but with some extensions that allow the provider "OAM"(operation, administration and management) capabilities. This is a layer 2 device.
They mention at the top because they didn't want to colocate it. So I imagine they did it because they could?
If I was filthy rich (which I’m not and probably never will be), I’d build a data centre in my mansion and buy a brand new IBM mainframe and stick it in there. Because I can. For the same reason other rich people build some huge garage with dozens of cars in it.
Mainstream platforms have what IBM calls FBA (Fixed Block Array) disks, in which the disk is an array of sectors which all have the same size. While the IBM mainframe hardware supports that, most IBM mainframe operating systems don't support those, however; Linux and VSEn are the exceptions, and z/VM is a partial exception.
As well as FBA, IBM has ECKD (Extended Count Key Data) disks, like the IBM 3390 was. (ECKD was preceded by CKD, which is conceptually the same – the same sector format, etc – but had a less efficient command set.) With ECKD, sectors can have variable sizes – each disk track can have different sized sectors, you can even mix different sector sizes within the same disk track. Also, sectors optionally have keys, and the disk has commands to search for sectors based on their keys. Originally, these variable sized sectors actually physically existed on disk; nowadays, the SAN only has standard hard disks (or SSDs), but ECKD disks are emulated in the SAN software on top of them. z/OS, the most popular IBM mainframe OS, only supports ECKD disks, not FBA ones. Linux and VSEn support both ECKD and FBA. For ECKD disks, you can't use the standard SCSI command set, you have to use IBM's DASD command set.
Many high-end enterprise arrays do support ECKD and FICON, but you usually have to pay significant extra license fees to enable the software that supports that.
In this particular case, there is no contemporary technological advantage of ECKD over standard hard disks/SSDs - it is simply a matter of legacy backward compatibility, lock-in, and an attempt to financially sustain the mainframe storage ecosystem. The idea of variable sector sizes may have had some real advantages when it was invented back in the 1960s, but nowadays it is just adding unnecessary complexity for no real benefit.
Even the idea of having the hard disk do key searches is something IBM has been moving away from, because doing them in software on the CPU turned out to be much faster in practice. It is still required because the primary traditional mainframe filesystem (VTOC) is based on it, but in areas where performance is critical (such as databases) IBM now positions it as a legacy technology.
Maybe there's still some value in moving key searches into the storage layer, but if there is, it would have to be something much more advanced than the rather simplistic 1960s implementation that ECKD storage provides. Oracle came up with a conceptually similar idea much more recently, Oracle Exadata (released 2008), in which some aspects of database query execution are delegated to the storage–especially table scans. ECKD only supports very basic comparison operations on the key field; Exadata, from what I understand (I used to work for Oracle, but not in this area) can evaluate complex predicates over the whole database row.
Samsung has a SmartSSD with a FPGA on it to offload compute to the drives.
https://phisonblog.com/why-is-computational-storage-inevitab...
>(as of this writing, VZ was pricing 50Mbps commited @ $455/mo; 100Mbps @ $661/mo; 1Gbps @ $999/mo; 5Gbps @ $2,099/mo; and 10Gbps @ $3,099/mo).
Ten gigabit is not yet a common residential internet service in North America. If that's what you want, this would be the only way to get it.
I've occasionally daydreamed about buying a condominium near one of the Seattle Internet Exchange's PoPs, buying a 400 gigabit port, paying to run a fiber pair, arranging for transit, etc etc. https://www.seattleix.net/join As expensive and pointless as buying a yacht, but at least getting it online would require a lot of tedious work!
(Like OP, I also jumped through the hoops for a Comcast fiber drop years ago (but to a business location) for the same eyewatering price but dramatically fewer megabits per second. The only upside is that the pair traveled directly to a CO in the same office park, so we would routinely see 1-2ms pings to anycast IPs like 4.2.2.2, 8.8.8.8, etc)
10Gbps PON is available from AT&T Fiber in some markets, and lower tiers (5Gbps) are coming online from Google Fiber in other markets. In SoCal, I've seen billboards for 10Gbps AT&T Fiber, for example. Obviously you don't get an SLA with a residential connection, though, and I highly doubt that many customers operating at full capacity will be problem-free.
For what it's worth, this isn't the same as what the person who wrote the post did. They bought a DIA circuit to an ISP.
What you're seeking is a point-to-point fiber link or MPLS to a panel in KOMO Plaza or the Westin building. You don't need to be near either of those locations to do it and CenturyLink (Lumen, these days) will sell you the circuit but you're not going to like what it costs. We have three of these circuits where I work and they are juuuuust a bit higher than what the article's author says they're paying for the DIA.
(This, along with the eyewatering cross-connect fees that the operators of those two facilities charge, is why getting a SIX port is so expensive on retail colo. You're either paying for the privilege of running some strands of glass or you're paying to use someone's extension switch and the connection they're paying for to get back to the SIX core.)
Given I have equipment already in the datacenter, it's pretty tempting to pull the trigger on a 100G link. I just have no idea what I'd do with it.
Completely agreed. Seattle and Portland seem to be high on the list of "places telcos think they can take complete advantage of." For what my employer pays for two (locked) racks of colocation here, we could get an entire small suite somewhere like LA or Chicago.
Wireline costs are similarly sky high, I think because CenturyLink has all of the cabling and no one has bothered to dig up the streets to run any more. When we connected a building in Eastlake to our old colo in Tukwila, CenturyLink was the only one who would quote us.
Those are the pings I see on my residential connection, though?
You get proper service: no random dropouts, and a clueful human on the phone and you could complain about QoS and they'd not laugh in your face.
The different thing here is that the service was provided to a residential building, but that's not entirely unheard of. I've done it myself for example.
In the event that there are issues at the primary location, internet routing protocols will re-route the relevant network prefixes to the secondary location.
The public /24 that was previously routed to their NYC datacenter now goes to that Connecticut house.
To an ISP, it means that a customer could exeperience a shortage of service, but that a group of customers would have to simultaneously utilize the infrastructure. ISPs try to keep a ratio of customers to service where this isn't likely to happen, or for discount providers, that this happens only during peak times. Ultimately it is a balance of customer attrition.
To a customer, it sounds like they are being deprived of what they are paying for. (Residental customers agree to a best-effort service typically - the opposite of agreeing to a specific service level). Even with ample service capacity, it is often the first factor considered when any component in the aggregate network is failing or underperforming.
I sure do not miss being a residential ISP. But I carry with me just enough sympathy for service providers that I am "always a blast at a party".
The problem I have is the lack of transparency, obtuse terms and the huge amount of funds taken in by the bigger operations that instead of being transformed into moar infra, get turned into exec bonuses/lobbying to not have to build infra/marketing to convince people the infra that was built is the best that's possible, yes sirree Bob.
When all you have is a string, and tin cans, you do the best you can. When your execs are taking big Federal bux, and setting them on fire instead of increasing overall throughput...
The problem I have is the lack of transparency, obtuse terms and the huge amount of funds taken in by the bigger operations that instead of being transformed into moar infra, get turned into exec bonuses/lobbying to not have to build infra/marketing to convince people the infra that was built is the best that's possible, yes sirree Bob.
When all you have is a string, and tin cans, you do the best you can. When your execs are taking big Federal bux, and setting them on fire instead of increasing overall throughput...
That's where I draw the line.
At home I pay for 5 but just need 2. I have failover to comcast. Works well . And yes I do get reported outages on the att side but because failover works I’ve not chased things down.
Of all of that the bulk of the price increase is them running the dedicated fiber so you can go from 32:1 to 1:1. It costs a lot to not have the same physical setup as the rest of your neighbors.
The monitoring and SLAs alone are probably the most significant differentiators, with these kinds of services you would usually expect an engineer on site within 4 hours to fix problems, and service credits for anything breaching SLA.
I once accidentally snagged a single mode fibre pigtail while moving stuff in a rack. Everything back up after the provider engineers had re-spliced a new pigtail on in less than two hours.
And that's before you consider the dynamic routing side of things.
(And the 3 x /24's that AS54316 have are also not cheap these days, that would cost something like $30-40k these days).