Showing posts with label Software Process. Show all posts
Showing posts with label Software Process. Show all posts

Wednesday, May 9, 2007

Worth Reading

While I'm working on my next big post, here are a few things worth reading.

Scott Rosenberg's post on ambiguity was right on. Ambiguity is a double edged sword - it can make things elegant, or intractable. Scott's insight is very sharp, as usual.

Last year, Basil Vandegriend put out a concise and helpful post on writing good unit tests. Most people agree tests are important, but many do not know precisely how to make them work. Basil addresses real issues, and gives good advice. I wish I had read this ten years ago.

Basil's latest post on the top five essential practices for writing software is also bang on. It is a quick must read for programmer trying to make the leap from just coding to professional software development.
Read more...

Saturday, April 7, 2007

Code Read 8 - Eric S Raymond's Cathedral and Bazaar

The ever evolving "The Cathedral and the Bazaar" by Eric S Raymond (aka CatB) has become something of a lengthy read. Scott Rosenberg's Code Read 8 dives on in, and rightly describes it as a "classic essay" and says it "has proved its importance" in the literature of software development. I first read it several years ago, it was considerably pithier then. However, it is still full of important ideas.

Cathedrals and Bazaars

The most important idea is its title track - two different ways to build software. In the "cathedral" style, a master architect and a small group of hand-picked skilled craftsmen work toward a grand vision with lots of direction and coordination. This is the traditional model of software development. The "bazaar" model, which is common to many open-source projects, particularly Linux, is one in which there are no hand-selected craftsmen, and no master plan. Instead the people in power select the best contributions from those people interested and able enough to contribute something worthwhile. The direction in which the project evolves is determined by the availability of volunteers willing to push it in that direction, as well as an entity, often an elected committee, which acts as an editor.

CatB presents a strong case that high quality software can be produced quickly and efficiently using the bazaar approach. Actually, to many people who just came into the software world in the last five years or so, this seems self-evident - Linux, Apache, MySQL, PHP, and Mozilla are just a few of the high-quality, feature rich, and reliable projects built in the bazaar mode. They are part of the landscape of the Internet which many take for granted. Yet less than 10 years ago there were many thoughtful, intelligent people who had serious concerns about the utility of any of these. Today, of course, the only people who seriously argue that open source development can not produce good software are those who have a stake in commercial alternatives. So while CatB's argument is convincing, it has already been won.

Much of CatB is also devoted to describing how successful open source projects work. If you are planning on running an open source project, these are must-read sections. But even if you are not, the idea of project leader as editor instead of architect is a very interesting one, and reflects a common gem of an idea in leadership theory that is subtle and often misunderstood: some of the best leaders do not lead by inspiring others to follow the leader's ideas, but rather the leader finds and supports the best ideas of the people they are leading. What makes Linus Torvalds (the creator and benevolent-dictator-for-life of Linux) such a genius is not his ability to write code or convince others of the correctness of his ideas (which both are probably impressive), but rather his ability to pick and choose the best from among the many contributions to Linux. There is a great deal that goes along with that style of leadership that is difficult for a type-A ego, and CatB delves into all of it with great insight.

The Twainian Passing of Closed Source Software

The second main thrust of CatB is that traditional management styles and closed-source software will ultimately be washed away under the coming wave of high-quality open source software and its management processes. This is a bit more controversial - and in some cases I think it is just plain wrong. Certainly it is possible that Apache may become the only web server anybody uses. It is also possible, although a bit more of a stretch, that open-source databases will replace commercial databases, or that Linux will become the dominant operating system, or that OpenOffice will become the only office productivity suite. But it unlikely that anybody but E-Bay will ever see the software that runs eBay, or that Google will ever open-source their search software. Sometimes, software is so tied to the fundamental service a business provides that there is just not enough interest in an open-source equivalent. We would all like to be making money like eBay, but how many of us are actually trying to write on-line auction software to match eBay's? There are simply to few developers to support it - especially since any E-Bay competitors needs to distinguish itself, which will probably require substantially different software. So there are some markets in which there is simply not enough demand for open-source software to make it viable.

Another example is corporate web sites, which will always be paid for in the traditional sense, even if they are built using entirely open-source software, simply because nobody but Acme Widgets needs an Acme Widgets web site.

This is one of the things I think Cat B misses in its open-source evangelism. Open-source projects work well when many programmers need the same thing to support the businesses they work for - so we see web servers, operating systems, programming languages and tools, web-site management software, a shopping cart or two, image and photo editing software, an office productivity suite, and so on. But the success of an open-source project depends on having a legion of programmers who need it. Yet CatB argues that commercial software is going away, along with all of the management processes that come with it.

I just do not buy it. Instead, I see the comercial software market becoming smaller, and more individualized. Nobody buys compilers anymore. GCC, gmake, Ant, Eclipise, and dozens of other free products all work fine. Moreover, support for open source products is often better than support for commercial equivalents. (I speak from personal experience.) A lot of software is going open source. But there will always be a market for software to support the specific business processes of individual organizations, which will require one-off construction - even if it does use off-the-shelf open-source components.

The Corporate Embrace of Open-Source

Indeed, many open-source projects are now significantly supported by companies (like IBM, RedHat, and others) that make money building exactly that kind of custom software. The underlying components are no longer something that distinguishes one competitor from another in the marketplace, so companies that compete against each other are joining forces to make everyone's job easier. Companies like IBM, RedHat, Oracle, and countless others are paying programmers to write code which the company will then give to its competitors.

This is one of the most interesting facets of the open-source revolution - the evolution of business strategy and corporate intellectual property policy in the face of the commodification of software infrastructure. In some ways, the software market has become an interesting experiment in altruism. More and more technology companies are realizing that despite superficial appearances, their software is not the core value they provide. And in that case, it makes financial sense to cooperate with their competitors (and anyone else who is interested) to build and maintain that software, while they focus on the what real value that they do provide.

P.S.

One final note - people can be slow to accept change. Some managers and programmers are still protective of "their" code against people in their own organization. This seems to me to be doubly backward, and I hope that as more and more open-source software succeeds, IT departments will learn let more open-source software practices through their cathedral doors. It can only make our lives easier.

Read more...

Friday, March 9, 2007

Code Read 7 - David Parnas on Star Wars Software

In 1985, David Parnas resigned from his position on a panel convened by the Strategic Defense Initiative Organization, which was overseeing the "Star Wars", or SDI anti-ballistic missile defense program. Along with his resignation he submitted several short essays explaining why he thought the software required for Star Wars could not be built, at that time or in the foreseeable future. Those essays were later collected and published, and are the subject of Code Read 7.

Some of the essays deal with issues such as using SDI to fund basic research (an idea in which he did not believe), or why AI would not solve the problems, but his core arguments focus around two main themes.

1) Software can not be reliable without extensive high-quality testing, and such testing could not be done for SDI.
2) Our ability to build software is insufficient to build SDI.

Scott Rosenberg, the author of Code Reads, seems to be asking what, if anything, has changed since 1985. Sadly, the answer is "Not much". Indeed, Parnas's paper is the most current of all the Code Reads sources, in is view of the software industry. It could have been written in 2007 just as easily as in 1985.

Testing is important

Of his two main themes, the first is the easiest to discuss. Basically, software is built broken. It needs to be tested before it works smoothly enough to be considered functional. Nobody has ever done it otherwise, despite their best efforts. Some have come close, but many have failed completely. All software needs to be refined, in situations very similar to its real usage, before it can be considered reliable. This is not news to anyone. Parnas makes it very clear how difficult it will be to do this with SDI.

Even if you exhaustively work to prove each component correct, or test each component extensively as you build it, the resulting system is still not trustworthy until it has been tested.

"If we wrote a formal specification for the software, we would have no way of proving that a program that satisfied the specification would actually do what we expected it to do. The specification itself might be wrong or incomplete." - David Parnas


A classic, and tragic, example of this problem is the Mars Climate Orbiter. Despite a rigorous testing process, a software error still caused the crash of the probe.

Software is hard

His arguments about our ability to build software go along three basic steps:

- Software is harder than other things we build.
- The way be build programs ensures there will be bugs.
- There does not seem to be hope for a much better way to build software.

Dijkstra, Brooks, and Knuth (as quoted in previous posts) have explained many reasons why software is hard. Parnas provides another reason - the discontinuity of software. He compares the structures of analog hardware, digital computer hardware, and software, and argues that since software is discontinuous, and has a large number of discrete states, it is much less amenable to mathematical analysis. This analysis is the main reason why non-software engineering projects are reliable.

For example, a structural member in a bridge has two states, "intact" and "failed". The behavior of the "intact" state is well understood: we have good mathematical models for the part's deformation under load, response to temperature, resistance to wind or water, degradation over time, and so on. The transition between "intact" and "failed" happens under fairly well understood circumstances. And we mostly just hope the "failed" state never happens. The same pattern of logic can be applied to almost all parts of the bridge.

Software systems, on the other hand, have many components, each with generally poorly understood behavior (compared to physical engineering), and many states. Indeed, most software approaches the what we now call "chaotic" behavior. Although it does always fulfill all three formal requirements, most software comes quite close. So on top of the layers of complexity, and depth of scale, most software is also, for practical purposes, chaotic.

How we build software with bugs

We try to manage this complexity by creating a logical model which we can use to break out smaller components, which themselves are broken into smaller components, and so on, until we are writing step-by-step instructions.

But this process is hard to do well. While we can write precise formal specs, "it is hard to make the decisions that must be made to write such a document. We often do not know how to make those decisions until we can play with the system... The result will be a structure that does not fully separate concerns and minimize complexity."

And "even in highly structured systems, surprises and unreliability occur because the human mind is not able to fully comprehend the many conditions that can arise because of the interaction of these components. Moreover, finding the right structure has proved to be very difficult. Well-structured real software systems are rare."

Additionally, we have the difficulty of translating those structures into code. Generally, we write programs as step-by-step algorithms, "thinking like a computer". We can sometimes do this in a top-down fashion, as Dijkstra proposed in his "Notes on Structured Programming", but even that uses a "do-this-then-do-that" approach. Various attempts have been made to find other ways, but none has found wide success.

"In recent years many programmers have tried to improve their working methods using a variety of software design approaches. However, when they get down to writing executable programs, they revert to the conventional way of thinking. I have yet to find a substantial program in practical use whose structure was not based on the expected execution sequence."

This provoked a heated discussion at Code Reads, but I think the fundamental point is that while other techniques exist and do provide real benefit in many cases, they are all ways of working with a larger problem, of structuring the overall approach. At its finest level, software is algorithmic, and algorithms are specified in sequential steps.

There are generally two main non-algorithmic ways to program. The first is to relieve ourselves of some of the work of creating complex sequential algorithms by specifying rules and having some system to implement those rules. And the second is formally isolating non-related portions of an algorithm so they can be run in parallel.

In the first case, using rules, is simple enough in restricted cases, but such systems become as complex and general programming languages in more general cases, and one eventually finds oneself creating rules to describe an underlying algorithm.

The second case works quite well, but each of the many parallel computations ends up being done using the same old step-by-step sequential executions.

And in both cases, we continue to make the same fundamental mistakes about the overall structure of our system because we don't fully understand its behavior.

Better tools and tecniques

Lastly, Parnas addresses the hope that improvements in methodology or tools will alleviate these problems. At the time he was writing, he saw four main threads in tool and methodology improvements (I'm combining two of his essays):

1) Structured programming
2) Formal abstraction
3) Cooperating sequential processes
4) Better housekeeping tools

Back in 1970's, according to Parnas, this was academic "motherhood" - nobody could object. Today, in my experience, this view is industry wide. A few people will argue that we've all been brainwashed and are now blind to alternatives, but even in the portion of our community which is most open to new ideas, these four still are the dominant paradigms.

Parnas argues that we are now in the days of incremental improvements in software, rather than rapid and dramatic advances, saying "Programing languages are now sufficiently flexible that we can use almost any of them for almost any task." And even things like non-algorithmic specifications still suffer from the same problems as writing code: "...our experience in writing nonalgorithmic specifications has shown that people make mistakes in writing them just as they do in writing algorithms. "

Conclusion

One could argue that Parnas is a dead-end thinker, saying nothing more than the status quo is bad, and it is all we will ever get. Instead, we must remember Parnas is talking about the most ambitious and complex software project ever conceived, and saying that the particular project is beyond our capabilities, not that software in general is beyond our capabilities.

However, I do think he misses a possible way out. I say "possible" because I do not know if it is a real solution, or just a fantasy. I think we need to improve the way we think about our own solutions. Then we can build systems that are less prone to the kinds of complexities which befuddle us.

To use an analogy, after the Tacoma Narrows Bridge disaster, civil engineers added some new factors to the way they think about bridges: "wind resistance" and "harmonic effects". Software engineers are still looking for what those factors are. I do not believe that we have a mature set of factors to consider when we design software, and I do believe that we can discover what those factors are.

In fact, I hope this blog will help me discover them, and I'd love to hear any suggestions.

Read more...

© 2007 Andrew Sacamano. All rights reserved.