I'd like Caltrain to publish raw train data
stackallocated.com
stackallocated.com
I did something similar using a GoPro and computer vision when I lived next to one of the 101 off-ramps in SF. I got it to work for most daylight hours (headlights screwed it up hardcore) before our landlord raised our rent by $1000/mo and we moved.
I figured it could have been a way to calculate ad impressions for billboards, but I also figured Clear Channel probably already knows those numbers.
Do you have industry experience with billboard advertising? I might rebuild it if there is an actual demand for it.
My solution, since I didn't want to do any screen scraping or make trying to identify individual busses/trains a project in and of itself, was to use Portland's TriMet API. That API acutally return specific route numbers, and estimated and scheduled times for each stop (interpolated in the case of non time points). I'm originally from the Portland Area, so I'm pretty familiar with the geography and roads.
From what I remember in the 511.org Google developer group, people have raised this exact issue, i.e. Caltrain train numbers. The guy responding from the MTA said they'd try and integrate it in the future, but these posts were like back in 2012 (IIRC).
NextBus is the source of bus position and predicted arrival times for MUNI, and appears to be the same for many other transit agencies. I can verify that (as of three minutes ago) it's still returning reasonable data.
However, if you're looking for the SFMTA schedule [0], I don't think you can get it through the NextBus API. I do know that you can get it through a GTFS "feed" found here: http://sfmta.com/about-sfmta/reports/gtfs-transit-data
Also, you might be interested in this, if you haven't seen it already: http://bdon.org/transit/ (SF MUNI transit delays. [This isn't my work.])
[0] Why would you want MUNI's schedule? It's not like any of the drivers care about it! ;)
Why won't it match the real schedule?
Caltrain also runs extra trains for special events such as baseball games. Those are planned, but not on the regular schedule.
Totally doable I should say.
- Eliminate at-grade crossings
- Switch to level platform boarding
- Double the number of trains
- Replace the current trains with faster accellerating electric ones
That ought to be done within a decade or two. Seems like more work than releasing the data they already have.Sounds like doable in a few years. You may even keep some at-grade crossings.
Whether or not level boarding happens, and what sort of service increases come with electrification are unclear.
Mentioned Moscow rail has lots of freight traffic and high-level boarding.
If caltrain adopted higher platforms, they would interfere with the size of freight trains passing through.
There are ways around this: If all the high platforms were on 4-tracked sections, the freight could use the express tracks only. Mixing high and low boarding might be more painful though.
Not to mention the current train stock doesn't support level boarding, so Caltrain would be restricted to use them as express trains and only have high platforms at local stations, for example.
Or the platforms could be far away from the trains and they use extending platforms: That is potentially unreliable and costly.
There is a somewhat reasonable blog about this type of issues at http://caltrain-hsr.blogspot.com/ which is mostly reasonable though it tends to be written as if it is the only reasonable choice.
I ride Caltrain every day and the experience annoys me a lot (still less than driving), but I feel sorry for the people running Caltrain because despite providing a valuable infrastructure service that is in huge demand, they are considered the lowest priority by everyone they interact with and have to deal with the most ridiculous restrictions and regulations. For example, Caltrain knows that they are always behind schedule during rush hour and want to adjust their schedule. But to do that they have to consult everyone and their uncle over a year-long process where every nutjob's concerns about a five minute scheduling change can stall the entire process.
Infrastructure projects deal with such problems everywhere, but I have not seen a place where this attitude is so deeply engrained and systematic like here.
By the end of the day, you're going to be more than a few minutes off from how you started.
Nonstop long-distance rail service can exactly match its schedule because there are relatively few places for entropy to creep in - you just hold a constant speed across miles and miles of track that you pretty much have to yourself. A commuter rail system is much more complex and there is much more room for entropy.
Of course, many American systems are operating at a pretty severe disadvantage, being hamstrung by poor equipment and infrastructure, a lack of funding, understaffing, a hostile political environment, and even pervasive cultural attitudes that dismiss railroads as being something worth investing in. I suppose given all that, it's a wonder they do as well as they do...
But still, I think it's important to never forget: it's absolutely possible to do much better.
You don't use CTA (Chicago transit) based on schedules; you go to your stop, read the board with accurate realtime predictions, and wait for a train to show up. It doesn't matter whether the schedule has anything to do with reality, just that you don't have to wait too long.
If you'd prefer not to spend too much time on the platform, then you can check your phone, which is accessing the realtime prediction feed anyway, to figure out when you should head to the station.
No part of normal usage of CTA depends in any way on the preset schedule, so spending money to keep to the schedule would be waste. Not only is it a Hard Problem, it's one that can be worked around very easily by providing realtime train location and prediction data.
Obviously frequent service is good, but there's a large variety of lines and sometimes very frequent service isn't warranted (or isn't possible because of funding/politics/etc).
For a line with non-frequent service, "just look at your smartphone" isn't a good solution, as (1) not everybody has a smartphone available, so it's a bad idea to make a transit service that requires one for decent service, and (2) more importantly, in many cases with less frequent lines not being able to plan can be a significant burden, especially when you need to make transfers along the way (in which case you always have to assume the worst case, and all the required uncertainty padding quickly adds up).
So for user convenience either you want frequent service, so planning isn't needed, or you want scheduling that's at least a little accurate, so planning is possible when necessary.
[There are also technical reasons for scheduling even very frequent trains in some cases, because when trains become frequent enough, the line itself can become a bottleneck, especially with complex services. For instance on many lines in Tokyo, both ordinary, express, and limited-express trains share the same tracks, with the expresses using in-station bypass tracks to overtake the non-expresses. The track network is also in many cases very non-linear, with many different lines sharing portions of tracks in some places, and different operators running onto each others' tracks. To do all this with high frequency services requires a delicate timing dance, and if you screw up the timing beyond some point, the whole thing very quickly falls apart.]
I've never ridden CTA, but this is the attitude that I treat Boston's T subway system with. It's wonderful, and like you say, you just don't care. Go to platform, board a train. Time-to-trains are acceptably low that you can go an wait. (looking at a random T schedule, if you arrive on the platform randomly, it's ~2-5 minutes wait on average if things are on schedule.)
That's not CalTrain. Trains are not as frequent (they I transit daily from Mountain View to SF: during the morning rush from 7-9am, trains are anywhere from 7 minutes to 34 minutes apart.) Missing a train also doesn't just mean the time waiting for the next train, but also lost time due to the train itself. (For example, the 8:05 is 7 minutes behind the 7:57 in Mountain View, but 15 behind in SF.)
Exceptional delays are significant, and not that exceptional. Trains hit things, catch up to other trains and follow them at a snail's pace, break down, stop for blocked tracks, can't be boarded due to overcrowding, departed early due to being "full", or are just missing without cause.
That said, if there was good data that just told me when trains would leave and when they would arrive, that'd be nice. I don't know of such a thing, and given the article, what exists looks clunky, complicated, and incorrect.
> Why should a transportation authority move heaven and earth to satisfy some moralistic concern about keeping exactly to schedules?
Because you're wasting people's time?
Another crowdsourced caltrain twitter account is https://twitter.com/caltrain You can see some of the more granular delays there. All these crowdsourced status accounts should be proof that caltrain SHOULD publish the raw data for us to use.
RE: scraping - instead of putting logic in your scraper, just download the entire section you need, store it in file format. Then parse and shove into database whenever you feel like it. You could rerun the parsing since you'll have all the historically scraped website data on disk.
Their commerce system is not mobile friendly and is a pain for mobile users. It could be much more efficient.