Simplicity and Utility, or Why SOAP Lost
keithba.net
keithba.net
For all intents and purposes, SOAP is a formatted pipe connection. The only issues I ever ran into were related to things like Java incorrectly implementing RSA encryption or Microsoft adding so much complexity that nobody used (and did not require).
There are still several major banks/insurance companies out there whose payroll transaction systems run on the SOAP implementations I put in place 10 or more years ago.
Now I see people claiming REST is better (and in many ways, they are correct), but I also see that large numbers of programmers don't really grasp REST any better than they did SOAP.
Did you write parsers for the messages by hand every time?
I did a bit of work with SOAP, and found interoperability to be a problem. Encapsulation and signatures both spring to mind. But then, i was using tooling on both sides. I shudder to think what dealing with SOAP messages manually would have been like.
Microsoft's 'WS' implementations were a mess of SOAP, but that's a whole different story.
a web classic, can never be linked enough times
Off topic question here. Does an archive.org article get cached so that when it's linked to it comes up faster? Or does it display at the same speed of a random archive.org page?
That page came up pretty fast for me don't know if that is random or because other people are hitting the link from HN.
< X-Page-Cache: HIT
[1] https://software.intel.com/sites/landingpage/pintool/docs/67...
I guess it's not fair to say that XED and DynASM are "much more complicated than the problems they're trying to solve." They are much more complicated than the problem I'm trying to solve. But I am surprised that there is no minimal X86 encoder with nice C++ syntax out there.
Also, another major issue with SOAP is that most all of the popular tools would generate classes based on a WSDL. This creates a toolchain issue that is readily solved for someone experienced, but can really suck if you are new to the idea of generated code working it's way into your project.
Much worse was the scenario of WSDL versioning in conjunction with these tools that generated classes. If the API never broke backwards compatibility, you should be OK and can just use the newer class representations, counting on the service and tools to deal with null fields appropriately (not always true unfortunately). But if version bumps of an API/WSDL broke backwards compatibility... then you had to have separate class hierarchies for different versions of the WSDL; what an intense headache. Contrast that to web REST APIs; without a formal schema and a lack of class-centric tooling, client libraries would often let the author stuff in the params themselves and let the serializer stuff in the body based on the params; so yes your API is not as formal, but this isn't a big issue but offers infinite flexibility in dealing with one-off versioning issues or interop issues.
As someone else mentioned, some platforms couldn't interop with others (differences as it related to simply stuff like nullable fields or primitives existed all over; total nightmare), and some features of WSDLs didn't translate well to certain languages.
A great WSDL author would know to make their WSDL as simple as possible, because they had spent time working with various language toolchains and new the limitations out there, but that's asking way too much and not realistic for someone to have to spend all their time trying to understand all the ways someone in language XYZ might use their WSDL.
We wanted automagic seriaization and de-serialization and we wanted every language-based construct (linked lists, arrays of structures that contain arrays) to be supported. Different tool vendors built that differently, hence the interop problems.
In REST/JSON, we don't have an IDL. We don't have a JSON Schema (yet! thank god) We generate human-readable documentation containing examples, for the definition. And developers build to those examples. There's nothing - no runtime or static tool - that gets in the way of developers serializing their data into JSON, the way they need to.
I suspect the dynamic language communities preferred JSON/HTTP so strongly because it was so much easier to use: they didn't have big corporate backers releasing SOAP tooling, but it's trivial to parse JSON straight into structures that are more or less native to the language (arrays of objects in JavaScript, lists of dicts in Python, arrays of hashes in Ruby).
Java did have heavy-duty SOAP tooling, and didn't (and doesn't) have a natural way to represent JSON's objects, so it was much less of a win there.
{"cities": [{"name": "London"}, {"name": "New York"}]}
Parsed into a variable called 'root' with either 'new JSONObject(...)' or 'json.loads(...)', compare: root.getJSONArray("cities").getJSONObject(0).getString("name");
To: root['cities'][0]['name']
The Python version is the sort of thing you find yourself writing in Python anyway (at least, if you write tabloid-level Python the way i do); it's reasonably idiomatic. The Java version is cumbersome and feels very un-Javaish.Bear in mind that with SOAP, in Java, by that point you would have mangled the data into realistic-looking objects, and would write something like:
root.getCities().get(0).getName();But has it really lost? I still see it used a lot in large corporate and high tech environments.
I was never a big fan of SOAP myself, but that's also partly because when I started, the tooling wasn't great and I got mostly a negative user-experience. But this is many years ago.
My understand is that these days it's quite good, and it has a pretty good user-experience. Furthermore, everything is backed by xml schemas, which makes it easier to create and require strictly valid xml being sent back and forward.
While some equivalents exist in REST-ish API's, it's arguably not very common and with SOAP it's a pretty much a free and built-in feature that everybody uses.
Modern web API design is great for getting things up and running fast, but it's not amazing for rigorously robust design.
I don't really like the "SOAP is bloated and over-engineered" narrative. It is uninteresting now, and was uninteresting 10 years ago.
I think it's much more interesting to learn about the things that SOAP did get right, and see how we can apply some of these learnings in our API's.
I'm writing something up on what was done right with SOAP and WS-*. (Arguably the code generation was useful for quite a few scenarios.)
Our mobile clients like REST because it's faster and easier to consume from outside the .NET platform than SOAP. Even within the .NET platform, a client can get up and running consuming a "RESTful" service in fewer lines of code than a SOAPy one. Not only SOAP tooling, but HTTP client tooling in general, is a lot easier to use these days.
WSDLs are cool because they are a fairly standard metadata format. Our clients like them because they can perform data validation with proxy/firewall systems before the request ever enters the corporate network. I've yet to see a concrete benefit from that kind of strictness, though. IMO, SOAP schema validation is so brittle that it causes more problems than it fixes.
Of course, you can also use non-SOAP document description formats like RAML for describing your RESTful services. It's not like SOAP offers anything you can't get for REST. But once you start adding all that on top of your REST API stack, you might as well use SOAP 1.1. (Never use SOAP 1.2).
I don't think the main problem with SOAP is that it's over-engineered, although it is. The issue is that the supposed benefit of "interoperability" was never really realized. It's supposedly protocol-agnostic, but nobody cares. It's supposedly interoperable, but everyone implements it differently.
What's wrong with SOAP 1.2?
If you are trying to communicate with a vendor that uses SOAP 1.2, it almost always seems to come down to guesswork. What set of properties do I need to specify before they accept my request? The WSDL gives you an object schema but it's not sufficient for the masses of WS-* extensions that you might have to support. Meanwhile, the vendor just exposes the WSDL and assumes that's sufficient "documentation" for clients and you don't need any examples or explanations or other info. I find the best way to approach this is to open Soap UI and start messing with settings until requests are accepted. There's no point asking the vendor for documentation because they probably don't know themselves.
Also, speaking from the perspective of a framework developer, I mainly start to pay attention when things stop working or clients are having problems. SOAP 1.2 doesn't have any more problems, but the problems are proportionally harder to solve. No matter how good your tooling is, it's not perfect, and it's a lot easier to look "under the covers" for RESTful services than for SOAP services. Of course that could also be attributed to the hellish nature of WCF, and not SOAP 1.2 per se.
...it almost always seems to come down to guesswork.
Meanwhile, the vendor just exposes the WSDL and assumes
that's sufficient "documentation" for clients and you
don't need any examples or explanations or other info. I
find the best way to approach this is to open Soap UI and
start messing with settings until requests are accepted.
There's no point asking the vendor for documentation
because they probably don't know themselves.
Having just dealt SOAP API last week, I can relate to this so much. I'm glad that I'm not the only one feeling the pain here.That's because of inertia -- a lot of those companies invested big in SOAP in the early to mid-2000s. They're not going to throw away an investment that big just because it's not cool anymore.
But don't mistake that inertia for SOAP being popular. It's just that so much money has been sunk into it that enterprise users are resigned to being stuck with it for a while.
XML nodes and children are just too slippery. I've actually argued with enterprise developers about why serving the following format was a terrible idea.
<user>
<id>123456</id>
<name>Joe</name>
<account>Account 1</account>
<account>Account 2</account>
</user> <user id="123456" name="Joe">
<account name="Account 1"/>
<account name="Account 1"/>
</user>I would have personally done
<user id="123456" name="Joe">
<account>Account 1</account>
<account>Account 2</account>
</user>
but I've never actually produced XML, only consumed.There is less flexibility there. Escaping seems harder, and you can't have sub-elements or lists. Avoiding attributes makes things more verbose, but then that's XML for you.
In your example, something like:
<users total="100" cursor="123453">
<user>
<id type="int">123456</id>
<name type="string">Joe</name>
<accounts>
<account>
<name type="string">Account 1</name>
</account>
</accounts>
</user>
</users>(Yes, you can wrap it with an <accounts> node and make your data extremely verbose.)
What about attributes? Should we ever expect them on the user? What if there's an attribute id as well, is it the same? (Oh, just download a schema to tell you what to expect! Argh...)
<user id="12345">
<name id="12345">Joe</name>
...
</user>
Let me count the ways that's underspecified in ways that can go wrong...(The actual problem I tried to solve was unrelated to this. The fields are in latin1 while the document itself is in utf8, with attributes specifying content type. It all validates apparently.)
Request:
GET /customers/43456 HTTP/1.1
Host: www.example.org
Response: HTTP/1.1 200 OK
Content-Type: text/xml; charset=utf-8
<customer>Foobar Quux, inc</customer>
Seems fine to me. In fact, having named elements without the requirement for a dictionary makes parsing straight to an object-representation (without an intermediate property-list) much easier.Using a simpler format for object serialization into XML is definitely an option, and it's a perfectly fine middle-ground between SOAPy verbosity and JSON compactness, but IMO it doesn't really have a lot to offer over a JSON version of the same API. Many API providers support both formats using content-type detection.
REST simplified things because the concept is about limiting the method invocation to the bare minimum and always have an explicit understanding of the state that is being changed or communicated.
The XML/JSON question is not as important as the method/resource question.
WCF and WS-* tried very hard to handle both RPC and more asynchronous scenarios, but the tooling continued to be heavily RPC based.
SOAP takes XML's problems to the next level.
When your manager tells you that your API needs to handle something "just in case", when your customer asks, "But what if I want to use this API with mumblefratz?", and when you look at your code and say, "Gee, if I make this a variable I could handle any special edge case by encapsulating its specialness in the variable ..."
It's clearly the result of developers and managers more interested in what the technology CAN do than what it should do.
http://msnpiki.msnfanatic.com/index.php/MSNP8:Getting_Detail...
and doing the same thing with the SOAP'd version:
>Here the WSDL and XML schemas for the web service descripted here, you can use them to generate a web service binding for your programming language.
>MSN Addressbook/Sharing Service WSDL & XSD files
This is how simple it actually is:
That link you gave has a summary part ( http://www.diveintopython.net/soap_web_services/summary.html ) that says "SOAP web services are very complicated."
With JSON, if the data returned by the server doesn't exactly match the documentation, it's the client's problem.
With SOAP, if the data returned by the server doesn't exactly match the documentation, the server is broken. The client, reading the WDSL file which documents the API, will report an error and refuse to continue. This means bug reports the server provider cannot make go away by the usual defect-denial techniques.
SOAP is past its peak but far from dead. In the Microsoft ecosystem, SOAP is nearly transparent, except for a bit of fiddling - you publish your WSDL from your definitions, I consume it to generate an API class on my end, and after maybe ten minutes on either side, we don't even think about SOAP again.
* Originally, Google Protocol Buffers. These are still in wide use; unfortunately, they never published/blessed an official RPC stack.
* Apache Thrift, aka Facebook's answer to Protocol Buffers.
* Capn Proto, by one of the Protocol Buffers authors, now has an RPC standard!
After having used Protocol Buffers extensively, I could never go back to untyped APIs. I'd at least use Apache Thrift for everything. Hopefully Capn Proto gets more languages supported for its RPC soon.
I'm still uncertain about melding the message format and RPC mechanics though. Part of me feels like network communications is just too big of a deal to abstract away into something that looks like a normal procedure call. But hey, even if that's always true for big systems, it probably won't be for smaller apps.
Being a Rails fan myself, I have always enjoyed using HTTP Web Services and interacting with tech companies' Restful APIs so easily.
This post provided a clear, side-by-side comparison of the technologies including the historical context of SOAP, which I found really useful and enlightening. If I could give you some Bitcoin for writing it, I would. Have a new Twitter follower instead.
It's hard to find succinct, objective technology articles in this day and age. Most reviews I find are either unconstructively biased ("dynamic typing sucks!") or too long and convoluted to be effective.
I never understood why it's out of favor?
A lot of vendors I've worked with will publish REST web services for public data but as soon as any private information needs to be accessed they use SOAP. I don't know if they just don't trust the tools to make REST secure or what but it's disappointing for sure.
Also the people who wrote the SOAP code are long gone which makes support for them even more difficult when I'm trying to get bugs fixed. REST server code has not been anywhere as bad to deal with.
1) JSON coming up to replace XML as a serialization technology, and
2) and Ruby on Rails (etc)
I remember even serious pressure from w/i the XML community in the sense of XML-RPC, and the joke that the entire spec for XML-RPC was smaller than the Table of Contents for the SOAP spec. SOAP was tough. It was confusing to use (fwiw, I had the best luck w/ Perls SOAP::Lite) and difficult enough to implement that devs I knew simply abandoned it.
Doing transactions and data format evolution in REST is hard, in SOAP it is simple.
The fact REST didn't provide some way for all methods to be invoked in a simple browser was to me a step towards the inaccessibility of SOAP.
http://api.rubyonrails.org/classes/ActionView/Helpers/FormTa...
That is, a POST containing _method=delete will be interpreted as a DELETE.
It seems faintly silly; if you're exposing an API to the browser, why use methods that you can't use directly and then have to go to the effort of emulating them, when you could just use supported methods directly? Is the use of the proper methods really that important?
But it does work.
And believe me - it was certainly not my choice.