Quality Culture – Part I

Quality isn’t just free … it’s profitable.

“A company with a highly developed ‘culture of quality’ spends, on average, $350M less annually fixing mistakes than a company with a poorly developed one.” HBRHarvard Business Review, in an article “Creating a Culture of Quality” from 2014, defined a “true culture of quality” as “an environment in which employees not only follow quality guidelines but also consistently see others taking quality-focused actions, hear others talking about quality, and feel quality all around them.”

The study was focused on big businesses – hence the $350M number – basically, it was discovered that between the bottom 5th and top 5th (quality culture), the top saved $13,400 annually per employee compared to the bottom.

The culture of quality includes four factors that drive quality as a cultural value:

1. Leadership emphasis; Managers are told that quality is a leadership priority. Managers “walk the talk” on quality. When evaluating employees, bosses emphasize the importance of quality.

2. Message Credibility; Messages are delivered by respected sources. Workers find that communications appeal to them personally. Messages are consistent and easy to understand.

3. Peer Involvement; Most employees have a strong network of peers for guidance. Peers routinely raise quality as a topic for team discussion. Like members of a sports team, peers hold one another accountable.

4. Employee Ownership; Workers clearly understand how quality fits with the job. Workers are empowered to make quality decisions. Workers are comfortable raising concerns about quality violations and challenging directives that detract from quality.

How do you know that your organization has weak culture of quality?

  • Lacks formal mechanisms for collecting and analyzing customer feedback
  • Experiences frequent, perhaps minor, setbacks owing to inconsistent quality
  • Training and development do no emphasize quality
  • Performance evaluation metrics have little/no mention of quality goals
  • Managers fail to consistently emphasize quality or resistant to quality initiatives

OK, fine for big businesses, but what does that mean for us, mere mortals? Small businesses? Stay tuned for part II!

Mistake Proofing

The initial term was “baka-yoke,” which means ‘fool-proofing’, but it was changed because the term ‘fool-proofing’ has dishonorable and offensive connotations. We Americans use a similar term: “idiot proofing” … which seems is even worse, but we don’t mind, do we?

Poka-Yoke (Japanese, pronounced PO-ka yo-KAY) which means ‘mistake-proofing’ or more literally – avoiding (yokeru) inadvertent errors (poka) – is the use of any automatic device or method that either (1) makes it impossible for an error to occur or (2) makes the error immediately obvious once it has occurred.

There are a couple of levels where this operates in software

In the SDLC – software development life cycle – we want the processes that are being executed to be as mistake-(idiot-)proof as possible … so we have high confidence that our product will be of high quality.

In the product itself … our requirements, design, implementation processes should ensure that the software product helps the user avoid mistakes.

When should we do this?

  1. when there is a process step where human error can cause mistakes or defects to occur, especially in processes that rely on the worker’s attention, skill, or experience – programming, code inspections, …
  2. At a hand-off step in a process, when output is transferred to another worker … programmer → QA → devops …
  3. Where minor early errors cause major problems later – e.g. getting the requirements wrong
  4. In the software product, where the user can make an error which affects the output
  5. In any case when the consequences of an error are expensive or dangerous.

Myths/Challenges

Because it originated as an industrial – a manufacturing – process, we can be led to believe that it doesn’t apply to software. After all, software is free to manufacture (reproduce) … these days we don’t even provide software on a physical medium (the youngest of you are saying “why would anybody do that?”), but poka-yoke is not restricted to manufacturing.

One software system with which I am extremely familiar, it is not only possible to make a mistake, but then have no way to recover – to fix the mistake – using the system. An expensive consultant needs to get into the database and manually (so to speak) adjust the data.

Even worse, you can get into a situation in the app where you are stuck and can’t get out except by force-quitting the app – we call that an oubliette.

A quick how-to:

Obtain or create a flowchart of the process. Review each step, thinking about where and when human errors are likely to occur. For a software product, use whatever design representation you have … you do have a design, don’t you?

For each potential error, work back through the process to find its source.

For each error, think of potential ways to make it impossible for the error to occur. Look for ways to:

  1. Elimination: eliminate the step that causes the error.
  2. Replacement: replace the step with an error-proof one.
  3. Facilitation: make the correct action far easier than the error.

If you cannot make it impossible for the error to occur, think of ways to detect the error and minimize its effects. Consider inspection methods, setting functions, and regulatory functions expanded on below.

Choose the best mistake-proofing method or device for each error. Test it, then implement it. Three kinds of inspection methods provide rapid feedback:

  1. Successive inspection is done at the next step of the process by the next worker.
  2. Self-inspection means workers check their own work immediately after doing it.
  3. Source inspection checks, before the process step takes place, that conditions are correct. Often it’s automatic and keeps the process from proceeding until conditions are right.

Setting functions are the methods by which a process parameter or product attribute is inspected for errors:

The contact or physical method checks a physical characteristic such as diameter or temperature, often using a sensor.

The motion-step or sequencing method checks the process sequence to make sure steps are done in order.

The fixed-value or grouping and counting method counts repetitions or parts, or it weighs an item to ensure completeness.

A fourth setting function is sometimes added, information enhancement, which makes sure information is available and perceivable when and where required.

Regulatory functions are signals that an error has occurred:

Warning functions are bells, buzzers, lights, and other sensory signals. Consider using color-coding, shapes, symbols, and distinctive sounds. In software we use pop-up warning or error windows (please provide relevant user-focused errors/warnings).

Control functions prevent the process from proceeding until the error is corrected (if the error has already taken place) or the conditions are correct (if the inspection was a source inspection and the error has not yet occurred).

#1 Challenge / thing to overcome

As usual, people. There is a certain amount of conscious reflection needed to apply this technique. Coding on autopilot perpetuates previous mistakes.

To recap …

Poka-Yoke technique is one of the highest value techniques in lean management. It is a way of ensuring quality without actually having a quality assurance process by preventing defects in the first place.

Poka-Yoke may be implemented in any industry – including software development – and has many benefits, viz.:

Helps work get done right the first time

Over time, it makes it more and more difficult for mistakes to occur

It’s “cheap as free” as strongbad would say …

Statistical Quality Control

There are only 7 basic tools needed to implement statistical quality control (and continual improvement); but why would you want to?

You may have heard of “Total Quality Management” or “TQM.”  If you have, you may think that its day has passed.  Well, it hasn’t … reports of its demise are … exaggerated. 

After WWII and even well into the 1960s, “Made in Japan” was synonymous with “cheap junk.”  In 1950, American Dr. W. Edwards Deming went to Japan to advise them on census issues and delivered a lecture to the Union of Japanese Scientists and Engineers (JUSE) on quality control.  Dr. Joseph Juran also delivered a lecture to JUSE on planning and management’s responsibility for quality.

Japanese industry adopted their philosophies wholeheartedly and 20, or so, years later (in the 1970s) the Japanese auto and electronics industries were “suddenly” crushing their American competitors.  Point of order here … Juran and Deming had been preaching their quality message to American industry for years before taking it to Japan.  American industry at that time was not ready to hear it, but getting hammered by the Japanese woke them up.

Consider that American industry had beaten the Germans and the Japanese in part by out producing them … American industry was, in a sense, still fighting the previous war.  But the peaceful war of trade was not being won by sheer volume.  But I digress …

Kaoru Ishikawa – professor of engineering at Tokyo University & father of “quality circles” – proposed the seven basic tools of quality control.  These have become known as the “seven basic tools” or the “seven old tools.”  There are seven “new tools” as well – these are management and planning tools which I will address at a later date.

The seven “old tools” will each be presented in a future post.  They are:

  • Pareto chart or diagram
  • Histogram
  • Scatter diagram
  • Stratification (sometimes replaced)
  • Shewhart’s control chart
  • Check sheet
  • Cause-and-effect diagram (a/k/a Ishikawa or fishbone chart)
  • In lieu of Stratification often either
    • Flow Chart
    • Run Chart

A word about data: collect it.  I guess that was 2 words.  “In God we trust. All others must bring data.” — W. Edwards Deming.  Collecting data can be tedious and the process may then be difficult to sell to team members – especially those who are “coding cowboys.”  The pushback will be that they have “real work” to do.  A bit of unsolicited advice: such people should be (VIPs) very important programmers … on somebody else’s project.  Your tools should facilitate the collection of data, but if they don’t then the folks need to …

To recap, after WWII, the Japanese took the quality control lessons of Americans like Deming, Juran, and Crosby to heart and went from being exporters of cheap junk to the makers of “the best stuff” – to quote Back to the Future.  Using the 7 basic tools of quality control, they rebuilt their economy and became synonymous with quality in manufacturing.  My upcoming mini course will introduce the 7 tools: Pareto chart or diagram, histogram, scatter diagram, stratification, control chart, check sheet, fishbone diagram, and I’ll toss in flow chart and run charts to cap things off.

Five Errors Managers Make About Quality

Often quality programs are not begun owing to errors in the thinking of management.  That might not surprise cynics, but it doesn’t have to be the case.

Lets clarify a few of the misconceptions about quality programs that might be lurking in the synapses of managers … if we are ever going to get a quality process in place

Crosby defines quality as “conformance to requirements” … the nice thing about that definition is that it is as meaningful for software as it is for manufactured products.  It also makes the definition quantitative rather than qualitative.  Because who wants a qualitative definition of quality?

Many people seem to think that quality has no objective standard.  That it is purely subjective.  The nice thing about “conformance to requirements” is that it makes it objective and therefore measurable.  And what you can measure, you can manage.

It is hard to “conform to requirements” that are undefined. We have worked for years with an ERP system that was clearly built without any overall vision of what it was trying to do.  You can always tell when software has been written without any defined set of requirements – and I don’t care what software development lifecycle model you use – if you fail to define what you are building before you build it then it will turn out to be a cobbled-together mish-mash of half-baked features.  

The 5 erroneous assumptions about quality plaguing management … and therefore organizations … prevent quality processes from getting a start.  Here they are:

Quality means “goodness” or “luxury” or … the relative worth standard.  Makes “quality” look like “gold-plating” and that reeks of expense with no ROI.  But taking the definition of “conformance to requirements” corrects that error.

Quality is intangible and therefore unmeasurable.  And what cannot be measured cannot be managed.  But it is measured by the cost – the expense – of non-conformance.  Re-work, fixing defects after they have been deployed, missed schedules and budgets all result from non-conformance.

There is an “economics of quality,” viz. an organization cannot afford to make things well.  Often this originates in a short-sighted or narrow point of view.  Fails to consider the overall lifecycle of the code – even if it is an MVP.

All problems of quality originate with the “worker bees” – especially in manufacturing (which doesn’t concern us in the software business) – in our case the programmers/coders.  In fact, quality processes should permeate the organization … the worst quality problems are those in management because they have the greatest multiplier.

Quality originates in (or is relegated to) the “quality department.”  Everybody in the organization is responsible for quality – has to care.

As Juran said: “It is most important that top management be quality-minded. In the absence of sincere manifestation of interest at the top, little will happen below.”  So top management needs to be convinced, persuaded … or the whole effort will go nowhere.

To recap … We define quality as “conformance to requirements,” and how closely a work product conforms to its requirements is measurable

That means as long as you take care to define the requirements then the quality will be measurable and therefore manageable.  And if you don’t define the requirements then just what the heck are you building?

You Are Not Your Code

The last thing I would ever suggest would be holding more meaningless meetings … the problem is not meetings, per se, rather the challenge is ensuring they are meaningful.

Once you learn to accept the fact that you are a fallible human being you will not get into a defensive crouch whenever somebody criticizes your work product (whatever it is).  Generally, your life is improved with the growth of any virtue – in this case, humility.  The big win is that your willing cooperation will help you become a better programmer/developer.

It is popularly – and incorrectly – believed that programmers have to be introverts.  An extroverted programmer is somebody who looks at your shoes when talking with you.  But programming, software development or engineering, is a people business .  Somebody who has trouble communicating with customers or other team members – is a liability – no matter how fast they can sling code.  In fact, they may be more of a liability the faster they sling code.

Recently, we were cleaning up some code that had been written by a former team mate and discovered that they had imported a bunch of complexity into a project to solve a fairly simple problem.  Two problems there … first, it indicated a massive failure of our code inspection process which we immediately moved to fix (that won’t happen again) … second,  the programmer was running open loop and had failed to consider what would happen if they weren’t the one maintaining the code.  Clearly, this could have been prevented by doing code reviews/inspections.

Not every organization is ready to implement inspections (a/k/a peer reviews).  A certain amount of preparation is needed.  Here is an inspection readiness framework to see if your team is ready. 

People

  • Champion
  • Management/leadership commitment
    • No commitment = lack of involvement, motivation
    • Do not proceed without upper level management commitment
  • Team member understanding
    • You are not your code – egoless programming
    • Personal commitment to quality above “being right”

Process

  • Policy structures
    • Review the code, not the person
    • Results not used for personnel measurement (may be used to direct remediation)
  • Training
    • Effective meetings (please!)
    • Inspection process

Technology

  • Tools – issue tracking, etc.
  • Measurement/metrics – defect rates/types
  • #1 Challenge / thing to overcome
    • Convincing management that a quality process – especially one that appears to be manpower-intensive – is crucial to the success of the product and, ultimately, the company.  Inspections, in particular, have a proven track record in the industry.
    • A decision of how much effort to spend on inspections comes from evaluating the risk-reward trade-off … time-to-market vs quality.  Note that even early adopters are unlikely to appreciate a product that produces wrong results or crashes often.

To review:

  • Inspections help us improve as programmers/developers
  • Even introverts can benefit from attending review meetings
  • We discovered – too late – that our inspection process failed and it caused rework reducing unneeded complexity introduced to a project by a now-departed programmer

Inspection readiness boils down to people – process – technology … each are important and the approach needs to be balanced

Convincing management to make a commitment to inspections will be the most significant challenge to overcome in getting an inspection process stood up.

Why Is Testing Not Quality Assurance

Dr. Edsger Dijkstra – one of the luminaries in the firmament of computer science – once said: “Program testing can be used to show the presence of bugs, but never to show their absence!”

The proper role of testing in the software development life cycle is to validate the soundness of the engineering process that produced the software.

Testing – as a component of the quality process – follows its own development process often, and properly, in parallel and coordinated with the development of the system.  That is how the “V” described in an earlier post functions irrespective of the actual defined development process in use.

Tests are planned and designed to provide coverage of the system – not only the externally visible interaction with the user, but also coverage of the structure of the code itself.  But that is not all there is.

What the system is supposed to do … usually called the system requirements … is established by the customer (internal or external).  This begins life as a more or less vague idea in the mind of the customer.  Any appreciably interesting system behavior is going to take more than one person to create.  Which means that things need to be written down.  Activities need to be coordinated.

For example, there is no way to “test” a statement of requirements … the only way to assure the quality (correctness and completeness) of the requirements is to review them with the customer.  So, the statement of requirements needs to be written to be communicated to the customer and then reviewed with them.

From the 30,000 ft level, none of the early effort in a project will be “testable.”  The tests themselves are not, strictly speaking, tested, either. 

  • How do you know that the test you are planning to run tests that part of the system that you were planning to test? 
  • How do you know that your test is not actually testing a completely different part of the system than the part you think you are testing?
  • What constitutes a correct set of pre-conditions (what must be true before running the test so that the test has validity) before executing a test?  How do you validate those?

None of this is meant to disparage software testing – it is a necessary activity in any development process.  The point is that testing is not all there is in a software quality process.

I can hear your unasked question: how much does all that cost?  How could it be afforded?  We will take up those questions in a later post.

How much refactoring is too much refactoring?

Here’s a question that came across the transom … it had to do with “… a 2-week Sprint where we went through a massive code refactoring in one of the major modules “ one of the quality engineers was wondering if the development team was being ineffective/inefficient …

The Question from reddit user

  • We are working in a 2-week Sprint where we went through a massive code refactoring in one of the major modules (a web module); the developers’ (5 of them) justifications were:
    1. At the time of the development, they didn’t know how to develop it
    2. To make the codes better for future development
  • After testing was completed, and suffice to say, was grueling as the tester had to go through multiple Sprints tickets and test cases just to re-test all the refactored codes, one of the developers took a ticket from the backlog where it also enhanced another “sub-part” of the module; therefore the tester now needed to re-test 30-50% from the same module again, and the test had to be done within the same Sprint.
    Unsurprising, the tester asked, “if this is the case why wasn’t that “improvement” done on the prior work since the work itself is very close to the already completed work? The developer said it will not be as difficult as the prior one, which seemed to be an attempt to appease the tester, but the tester was not having any of it and complained to the test manager that the development process is being ineffective and inefficient in their development work, wasting testing time.
  • To the experienced QAs, what do you think of this situation? Is the development team being ineffective and inefficient? Was the tester right in concluding as such? What went wrong?

My response:

  • The question posed (constant code refactoring) and the scenario presented looked like two different things. So, first step is to validate that we were dealing with a chronic – rather than an acute – problem.
  • If the “massive code refactoring” was needed – and both rationales presented appear to be adequate reason for re-factoring – then there should be little argument about doing it. That doesn’t make the refactoring “constant” unless it recurs sprint-after-sprint-after-sprint …
  • If there is constant refactoring then there is a serious project management/engineering breakdown. The quality problem in such a case is far removed from testing because the is a systemic process problem leads to a lack of stability in the product.
  • In the case where there is a single “massive re-factoring” (not constant) then … having left major issues lying around in the backlog and/or allowing major issues to be picked up late in the sprint … that is a different project/engineering management issue. It suggests that the estimates of the effort required may have been be seriously flawed.

What does it take – what does it cost – to write one line of code?  What do you include in the “what does it take” – documentation (user and developer), what quality assurance activities are needed, what tools or environments are needed, etc.

In our consulting practice, we have seen projects running with no estimates whatsoever.  And no process for improving estimates, if there were any.  How do you know how many user stories can be developed or issues resolved in a sprint if you have no idea of what it will take? 

Estimation of the complexity of software work is notoriously difficult.  New, green field development, often has so many unknowns that it seems impossible to estimate.  My very first comp sci instructor’s rule of thumb seems to work:

  • Take your first impression (guess) of the time it will take.
  • Multiply it by 2 (or 102)
  • Increase the units by 1 order of magnitude
  • So … an original estimate of 4 hours becomes … 8 days

It is appropriate for a manager (or customer) to ask: what will it take (cost, effort) to get something done?  How long – wall clock time – will it take?  If you cannot answer those questions then it doesn’t matter how “agile” your process might be … you should not proceed.

The Hero Syndrome

A “super programmer” can be 5x – 10x more productive than the average programmer. That is an astonishing difference and you will know if you ever work with one … it is obvious that they just “get it” better and faster than those around them.  Does that make them a good and valuable member of the team?  Well, maybe.

Patton once said “… a staff officer of uncongenial disposition should be relieved irrespective of any other qualities.”  Presumably including their productivity.  The point he was making was that the staff was a team and the friction caused by an ill-disposed person would be destructive to the team morale and overall team performance.

The seduction of hiring and retaining a hero is that things get done … and done fast.  The question, of course, is what gets done?  Is it the right thing that is getting done – the team’s vision of what is to be produced – or just the hero’s vision.  Tony Stark works alone … OK, with Jeeves the bot … basically alone.  He works on his vision of what is needed.

It is not impossible for a hero to be a good team player.  Take a virtuoso violinist … if the orchestra is playing Beethoven and the violinist is playing Mozart … well, he might be playing Mozart the best it has ever been played … and it will all sound like crap.

Keeping the hero on board with buy-in to the overall project and what is being developed by others can be a particular managerial challenge.  If your favorite hero can only see their own vision – rather than make the team’s vision their own – then you will have a challenge.  The hero may generate so much code so fast that it looks like waste to toss it out … easier to switch the vision.  The waste, though, was on the part of the hero … they wasted their effort on their vision rather than the team’s.  That is not to say that they mightn’t actually be correct and the team might really need to change their vision … 

Attrition.  Well, the hero rides out of town and the townsfolk are left with … themselves.  The more the organization has relied on the hero, the more difficult this is.  The team/organization has to be prepared to lose any team member … at any time … for any or no reason.  Murphy’s laws of software development includes:

  • The more important a resource is to completing development, the more likely it will be missing when needed most.
  • Corollary: If a resource is needed to complete development then it will go missing and do so at the least opportune moment.

The right way to employ a hero may be as a hired gun … a contractor … bring the hero in to solve a specific problem and then let them go.  The problem specification should include that the solution is documented so that the team can sustain it long-term.