Saturday, May 4, 2013

Structuring Software Engineering Case Studies to Cover Multiple Perspectives

This post is my riff on a paper titled Structuring Software Engineering Case Studies to Cover Multiple Perspectives by Emil Boerjesson and Rober Feldt from Chalmers Univ. The paper offers suggestions on how multiple perspectives can be ensured by using a 6 step process. In their case study, they wanted to ensure they looked at the V&V process using four different perspectives; Business, Architecture, Process and Organization (BAPO) as well as from three distinct temporal perspectives; past, present, future (PCF).

The paper does not have any deep contribution to the case study approach to software engineering research; however it does provide an easy paper for the start of understanding how case studies can be used in the research of software engineering.

They call a case study:

  • an observational research method,
  • a way to investigate a phenomenon in its context,
  • applicable when there is no clear distinction between the phenomena and its context,
  • have guidelines recently published by Runeson and Hoest [1]

Their six-step methodology is:
  1. Get Knowledge about the Domain
  2. Develop focus Questions/Areas
  3. Choice of Detailed Research Methods
  4. Data collection
  5. Data analysis and Alignment
  6. Valuation Discussion














[1] P. Runeson and M. Hoest, "Guidelines for conducting and reporting case study research in software engineering," Empirical software Engineering, vol. 14, no. 2, pp. 131-164, 2009.


The Accretion of Structure in Large, Dynamic Software Systems: A Socio-Technical View

On the outside chance there is really anyone following this blog, apologies. If you acknowledge reading this I will be more inclined to post more often. But this blog serves as my sounding board for my developing research thoughts. The title of this post could potentially be a dissertation but it is too ambitious to be mine. Yet it does indicate the direction my thoughts are taking.

Large software systems that reach any great level of popularity do not fall from the sky fully formed. They may have germinated from the seed of some idea that one individual had or they may have been created to solve some problem. Either way, that is only the origin story of the product. To reach its mature form, required the efforts of many people over some period of time and we classically refer to this as a project. A project has the creation of the product as its end goal. Once that goal is achieved the classical view is that some organization, an organization that may have been created coincident to the creation of the product, will take ownership of the product and use that asset for the betterment of the organization. If the single product is the sole asset of that organization, its fortune will be tied to the product life-cycle of that asset and will cease to exist with the retirement of that product.

Contemporary software engineering research owes a debt of gratitude to the open source movement. The benefit of the movement to research is the open nature of the work. Rather than guard the artifacts of software development behind the shroud of proprietary secrets, the open source community thrives on an atmosphere of complete transparency. Even better, it has a tradition of maintaining these records for posterity making longitudinal study possible in a way that is unthinkable with for-profit corporate development. The oldest and best known of these efforts is the Apache web server and the organization it spawned. It is this organization and the many products that can be tied to this organization that has been my latest object of research.

The history of the Apache Software Foundation is well known. It has matured as an organization over the years and epitomizes a form of software creation and stewardship that is neither completely altruistic and selfless nor constrained by the commercial marketplace for software. As such, it may, or may not, result in software that is qualitatively different over the product life-cycle than the products from such well known software houses as Microsoft and Oracle. But whether they are different or not, it is possible to study the inception, growth, decay and retirement of these software products in ways that are impossible for the commercial products. It is the forces of this evolution that I am drawn to and am looking very closely at in the various Apache projects. (Note that unlike a classic project, an Apache project is really an organization that is responsible for both the creation, maintenance and general stewardship of the software product until its retirement.)

A thesis I have is that program correctness is not really the most appropriate aspect of the dynamic code maintenance that most of my colleagues think it is. Rather I believe that it is the change in the non-functional behavior of the product that drives more decisions, and ultimately predicts the ultimate failure of the product. The froth that is seen around the correctness of the product is certainly important. But a product that fails to deliver in performance, scalability, maintainability, adaptability or any of dozens of other qualities will either find a niche or will need to be re-engineered to meet these new challenges.

I say re-engineered in the sense that the small focus of most enhancement requests or bug reports do not allow for a scope that is sufficient to address the refactoring that is most often needed to achieve these ends. In many cases the skills gained by  the creation of this product result in a functional team that are motivated to either significantly enhance the product to give it those qualities or to leave the product as is and create a new one via fork or greenfield development that will improve upon the prior product in these key qualities. Often the relationships between these distinct products is overlooked and the debt one product owes to a prior product lost in the mist of time.

The history of the Apache Software Foundation is celebrated. But every day there are projects that choose to put their products into retirement because of lack of interest. If other projects were spawned by the project, they are sporadically documented, sometimes in the project archives and communications, or sometimes through the press which covers open source development. What I have lately come to realize is that there is no interest in documenting this history and as web links rot or repositories taken off line, history is being lost every day. It is perhaps Quixotic to concern myself with trying to prevent this complete loss but I do. As such, I am trying to document as many origin stories for Apache projects as I can in the belief that if I do not, some of this information will eventually become unattainable. What is harder is to justify this effort in terms of my own research.

No major software product exists that did not require thousands of decisions, some big, most small, made by the team members over time. Software developers are notoriously averse to documentation so these decisions are as often inferred as observed. This is where Big Data and the statistical methods of empirical software engineering are most helpful. But where the data is not available in a form usable by these techniques, the techniques are useless. Since the gathering of this data and putting it into a form that is suitable for quantitative analysis is exceedingly tedious, it is usually not done. In this reticence to tackle difficult data gathering tasks, I see an opportunity. I see this as an opportunity to both do the deep dive into more interesting decisions made during software development, their relationship to the form and structure of the software system and also make a significant contribution to the preservation of some important history that may have value for researchers in the future. Perhaps I am suffering from a touch of hubris in this but it helps my motivation to feel that my contribution may surpass my own prosaic goals of satisfying institutional requirements for a degree or getting some papers published.

I'll end my post here since the work is ongoing and quite fluid. But this post gives me a short statement that I can share with others who may take an interest in my current direction.


Saturday, March 23, 2013

Quotes from The End of History and the Last Man by Francis Fukuyama. First Free Press trade paperback edition 2006

A seat mate of mine on a flight from DC to Sacramento once suggested that this book is an influential book for the neocon movement and hence it went on my list of books to read. What follows are quotes that I might be able to use in the future.

"But the truth is considerably more complicated, for the success of liberal politics and liberal economics frequently rests on irrational forms of recognition that liberalism was supposed to overcome. For democracy to work, citizens need to develop an irrational pride in their own democratic institutions, and must also develop what Tocqueville called the 'art of associating' which rests on the prideful attachment to small communities. These communities are frequently based on religion, ethnicity, or other forms of recognition that fall short of the universal recognition on which the liberal state is based. The same is true for liberal economics. Labor has traditionally been understood in the Western liberal economic tradition as an essentially unpleasant activity undertaken for the sake of the satisfaction of human desires and the relief of human pain. But in certain cultures with a strong work ethic, such as that of the Protestant entrepreneurs who created European capitalism, or of the elites who modernized Japan after the Meiji restoration, work was also undertaken for the sake of recognition. To this day, the work ethic in many Asian countries is sustained not so much by material incentives as by the recognition provided for work by overlapping social groups, from the family to the nation, on which these societies are based. This suggests that liberal economics succeeds not simply on the basis of liberal principles, but requires irrational forms of thymos as well."
pp xix-xx


Friday, January 25, 2013

Our paper this week is Echoes of Power: Language Effects and Power Differences in Social Interaction  which explores how to identify power differences between people in a domain independent way. http://www.mpi-sws.org/~cristian/Echoes_of_power_files/echoes_of_power.pdf.

The central thought is that the way in which one person "coordinates" their linguistic style to the style of the person or group with which they are communicating can be an indicator of the power relationship between them. If this is true for open source software as well as for the Wikipedia and Supreme Court corpa they explore, it can be exploited to create a power hierarchy in these open source projects. With a graph of the power network it may be possible to infer project outcomes based on various network metrics of that power network.

This paper also makes me wonder to what extent it could be extended to code style. The references suggest that they are building on some other work that finds a prose style that is characteristic for an individual. Given the high level of semantics in the tokens and rigid syntax in a computer program is it even possible to find some domain independent marker of programmer style? I know this has been extensively explored and this paper makes me more curious to see what has already been found.

Friday, August 10, 2012

Free Software and Communism

Sac State has a program through MSDN that provides their current OS for free to the students. UC Davis does not (or at least I haven't found it yet if they do). Curious. But the bigger issue is my changing thoughts about what I expect to spend for software.

In the 30 years I have had an association with Microsoft I have probably spent in excess of $10,000 buying their products. Despite this continued relationship Microsoft doesn't know I exist and provides some of the worst customer service I have ever experienced. I have had a few times where it was a bug in their product I wanted to talk about and have been given the run-around or been asked to fork over $100 to talk to someone. Microsoft has historically charged too much for an inferior product for far too long. They seem to relate to the consumer market the way a farmer relates to a field crop.

While I myself am not the sort inclined to tinker with an OS, many people I have known are. The Linux and open source movement have been incredibly empowering for those of us who have the capability and inclination to tinker. First it is free. That of course is a huge advantage seeing that Microsoft has the audacity to float an undiscounted price for Windows 7 that is over $500. If I hadn't been getting their products for free from school I would have abandoned them long ago. There is a certain hubris that the people they might want to recruit to write operating systems would be asked to spend a significant amount of money in their education using scarce resources to buy their products. It would be in their interest to ensure that every CS major receives a generous support from them if for no other reason than good PR. So at some level good CS students must find it galling that they are asked to pay a high price for an inferior product. Is it any wonder that it is often students are the most creative in finding ways to avoid paying for their products?

Since it is CS students and the professionals they become use software as their tools of the trade daily over their careers, it makes perfect sense that they will be the most critical consumers of these tools. Any worker who uses a tool daily will choose and maintain her tools with great care. When the market provides only a mediocre tool, there is an opportunity for some worker to turn their attention to the tool. Of course many workers do not have the skills to tinker with their tools. A machine shop worker probably does not have the technology at hand to tinker with a high-precision lathe. But a software engineer with the source code for a major piece of system's software does. While it is only a minority of these engineers who are inclined to tinker, it is enough to create an alternative in the tool market. We have certainly seen that exhibited in the market for open source software and the active support it gets.

In some idealized libertarian universe, these talented tinkerers could have formed an alternative to Microsoft and made some money from their collective efforts. But creating and managing a business is not a trivial affair and certainly not what motivates talented software engineers. Engineers in general are great admirers of functionality and not of management. Their are significant intrinsic rewards to the creation of a good product but few in the marketing and sales of that product to them. Yes, there are some rare individuals who excel at both. But their scarcity relative to the number of highly talented engineers is precisely my point. The open source model was appealing to them because it gave them control over their means of production and empowered them to do things for their clients that are impossible in a closed  model.

Some have tried to cast these engineers as proto-communists contributing to a common good. (from each according to his ability, to each according to their need). But if that attitude exists in the open source community, I have not yet seen it baldly stated. Rather than being driven by some ideological position, the market seems to be moving along a highly pragmatic path of least resistance. The closed software market has not met their needs to provide high-quality, responsive solution to the clients using their tools. Instead of forming highly responsive relationships with these integrators, the companies seem to have put their needs behind the company's. Doesn't it make sense that someone responsible for providing the company with a high-quality web server would be more inclined to support Apache than IIS? It is in their self-interest to choose this simply because they have greater control over the resolution of any problem or enhancement than they would over a close product at the leading edge of its deployment. The fact that other people gain access to a high quality product as a consequence is irrelevant to them and takes nothing from them. In fact, the widespread adoption of Apache enhances their marketability since employers who desire to avoid IIS, for whatever reason, will seek them for employment.

No, I am not finding any incipient communism in the growing open source movement but instead a marketplace that is responding to the maxim 'information wants to be free" plus an enlightened self-interest among the most talented software engineers of our day. They seem to be having a blast. Now if we can only figure out where the margins of this new market lie.

Wednesday, August 1, 2012

A First Look at Commit Data for an Open Source Software Project

I got a little something to sink my teeth into. I am looking at commit data for some open source projects. This is mostly an exercise in regaining my sql chops and learning R. Here is my first plot:


The plot has the alias of the submitter along the X and the timestamp along the Y. What you can see for this one project is that there are two heavily dedicated submitters who both started working on the project about the same time, one who is more sporadic and started shortly before the two of them and two who appear to have started the project but only periodically commit although their activity is relatively consistent over the entire length of the project. What is somewhat surprising is how many there are who have almost no commit activity (there is some doubt regarding whether this alias is the submitter or committer although it is supposed to be the commit id). It seems odd that someone would have gained committer status and then stopped commiting. Will I see this pattern in other projects? Who are these heavy committers and how do they differ in role from the people who appear to have started the project?

Tuesday, July 10, 2012

Will Open Source Software Someday Rule the World?

I've been listening to the ChangeLog and reading a lot about the explosive growth of open source software. There is no doubt that this phenomenon is creating innovative products and exploring new ways for creative individuals to self-organize. I forgive young software engineers for thinking this is a complete game changer and that all software in the future will be developed using this model. Perhaps they're right but I don't think so. I want to explore this assertion and see what the data shows. Let's start with the Apache Foundation.

In their how-it-works page, they describe how the Apache server source code was taken from the NCSA HTTPD but was orphaned when the original developers moved on to other projects. Brian Behlendorf created a mailing list of other admins using the product to communicate patches on that code base. That original group continued to organize and formalize. They then began to sponsor allied projects. All the projects were organized to grant authority to volunteers after they had proven their abilities; a culture of meritocracy. This group was not only users of the products, they also had the skills to modify the product and the reasons to do so. Their working lives revolved around the functions of these products and many changes were either self-serving or would please they people who they served.

These open source projects create a resource that is designed as non-excludable and non-rivalrous. Anyone who find the repository is welcome to download the software product and use it. The use of the product by one person does not preclude the use by anyone else. In many ways this is the classic public goods model. What is interesting is how the FLOSS projects do not seem to suffer from the free loader problem or any form of tragedy of the commons.

Someone who downloads a software product from a FLOSS (Free/Libre Open Source Software) repo will often have access to a significant asset. The Apache web server is a sufficient example. This product competes quite well against products from for-profit enterprises and could possibly attain a dominant position in the marketplace, if it hasn't already. But to achieve and maintain a market position like this required a significant investment in effort at one point. In the case of the Apache webserver it was NCSA. If the creation of a superior product that is then gifted to the public, this matter would be easily understood. The Apache webserver was not a single creation that was gifted since the needs of webservers have changed dramatically over the years. In fact it was the new needs that motivated the creation of the project in the first place. Since the product was already a reasonable base upon which to build, and since the needs affected many users of the product, a natural mutual interest was created and nurtured by the communications channels that were available. This is still the model for most people of a FLOSS project.

If the common asset were open to modification by anyone who had access to it, it would surely have been doomed by the tragedy of the commons. Each person who thought they knew what they were doing would have made changes and some of them would have broken the system. The form of ownership that evolved was that only a subset of people would have commit rights and they already had established enough trust in each other to be assured that they would act for the common good of the project. Since these "managers" depended upon the product, their interests converged.

This software, unlike physical public goods, could be consumed without limit as it existed. But to remain viable in the marketplace, it needed to evolve to meet the changing requirements, lest it become obsolete. Here is where the Apache project had a distinct advantage; each consumer also had the requisite tools and knowledge to adapt the product for the changing needs. These changes were insignificant compared to the complete body of code and the cost of creating each patch was also insignificant to the cost of changing to another product. Once the patch had been created, the marginal cost to share it with the community was insignificant and its value to the individual was limited unless that developer/user chose to take on the responsibility for creating a marketable product from the FLOSS base. Since this was not compatible with that person's core mission (we assume), they chose to contribute. Since each person was acting in their self-interest, the collective effort not only allowed the product to evolve to meet new demands but could also spawn new products as needs diverged.

So what are the elements to create a successful FLOSS project? First, there must be some base of code that provides enough out-of-the-box functionality to answer some need by a user. Second, there must be some core set of users who have a shared vision of what the product is and can agree on commits to that product. It seems as if this core group must initially be the primary source of commits as well as be willing to support new users of the product. Third, the end-users of the product must also possess the tools, skills and motivation to patch the product when it does not meet their needs and offer these patches back to the project. It is possible that the work of Ronald Coase may be helpful in the analysis of FLOSS projects.

The FLOSS market has changed over the years and it appears that other models are appearing. In particular the market for ERP software in FLOSS looks very different. In this case, the initial investment to create such a product is much higher and the market much thinner. What organizations now seem to try is to offer a "community version" of their product which is appealing to some subset of the market and then use it as a loss leader or teaser to find others who might be willing to purchase a traditional license for the full-product.

Another model I see in the repos is people who have come up with a relatively small widget which they have gifted. It is not clear if their hope is to spawn a new open source project from this seed or whether they are just engaging in a form of gifting. My instinct tells me the chances of a relatively small product spawning a viable new project is small. First, it may make more sense for the the consumer to merely copy the code into their own repository rather than take the product in as a component. If that is the case, the opportunity to achieve a back channel of patches to the product will be undermined and the product will remain static. If there are any changes in the needs, the product is likely to wither from the lack of contributions to keep it viable in the market. Second, without a core of developers who see value in providing the oversight to the project that is needed to avoid code rot, the burden will fall entirely upon the original developer. If that developer does not make frequent mods to the code herself, it is likely to become a burden since there will be a limit to the time that the developer will devote to that one product.

Eclipse seems to be following a different model where instead of individuals supporting the product it is organizations. This is an area I need to research since it is not clear to me how IBM benefits from the expense in Eclipse.