No More Than Two
fixed-length iterations: a transitional practice
Suppose that you had an iteration of one week, followed by an iteration of 2 days, followed by an iteration of 1 day, followed by an iteration of one-half a day, and so on. If you still had your sanity at the end of this process, would you have learned anything? I haven’t tried it with a team yet, but here’s the thing that I hope would come across: if you apply enough ingenuity and you’ve acquired enough skill, you can deliver business value in shorter times than you can currently imagine.That would be cool. Of course, the Kanbanista's seem to suggest going straight to that world in one step. And maybe that can work in a certain setting, and maybe not. The idea scares me, whe I look at most of the teams I help.
Government IT projects: who can politicians listen to?
- Accenture build a system for the RPA: not fit for purpose, £46.5 million overspend
- BT, Fujitsu and others build NPfIT for the NHS: not fit for purpose, £10.4 billion (with a “b”) overspend
- Fujitsu build an information system for magistrates courts: £342 million overspend
- Cap Gemini build PRISM for the FCO: not fit for purpose, £34.5 million overspend
[the Conservatives] began to take on some of our suggestions, as they came to better understand government IT. For example, their proposal to cut IT contracts into smaller and shorter chunks was dropped as they realised they would have to act as system integrator to each of these smaller projects.What is the taxpayer to do in the face of this sort of thing? Particularly the well–informed taxpayer who knows full well that in no way whatsoever is this argument from Mr Carter valid.
T-shaped designers
Tests vs checks
Software Engineering?
I can hope.
- Look at what engineers ‘do’, not what they build.
- Catch up with the state of the art in what is conventionally called engineering.
Innovation
The Shock of the Old
The Revolutionary Period of Big Innovation
in recent years it has started to look like we're moving out of the revolutionary period of big innovation, and into a phase of relative stability.I don't believe this for a minute.
Sources of Diffusion
no matter how good and powerful our software tools get, we are only getting a fraction of the leverage out of them that we could get.Programming tools are no longer where the greatest potential lies.
We will get the biggest leverage, not just in programming but in all our endeavors, by discovering better ways to work together. [emphasis in original]
We are uncovering better ways of developing
software by doing it and helping others do it.
Through this work we have come to value:
Individuals and interactions over processes and tools [etc]
We were doing incremented development as early as 1957, in Los Angeles, under the direction of Bernie Dimsdale. He was a colleague of John von Neumann, so perhaps he learned it there, or assumed it was totally natural. [...] the technique used was, as far as I can tell, indistinguishable from XP. [...] much of the same team was reassembled [...] in 1958 to develop Project Mercury, we had our own machine [...] whose symbolic modification and assembly allowed us to build the system incrementally, which we did, with great success.
UK Goverment IT Failures
[...]the total cost of Labour's 10 most notorious IT failures is equivalent to more than half of the budget for Britain's schools last year. Parliament's spending watchdog has described the projects as "fundamentally flawed" and blamed ministers for "stupendous incompetence" in managing them.
I have no doubt that ministers are, as Michael Savage puts it, "too easily wooed by suppliers". The suppliers which we see winning government contracts again and again might also find it too easy to woo a minister. Particularly a minister and a department which would rather launch a high-profile, high-budget, high-risk project than adopt the smaller scale, incremental approach that creates few headlines when launched, but also has a much better chance of creating few headlines when it does not fail.
It's the Screaming
The community of developers whose work you see on the Web, who probably don’t know what ADO or UML or JPA even stand for, deploy better systems at less cost in less time at lower risk than we see in the Enterprise. [emphasis in original]That's a pretty bold claim.
Scale
Screaming
Solutions?
dynamic languages and Web frameworks and TDD and REST and Open Source and NoSQL at varying levels of relative importanceToo right. But, as Gerry Weinberg says: it's always a people problem. It always is. It's the screaming.
Complex Domains: playing with Alloy
Microcell Predictor
Semi-formal
This is Now
mechanising such a large proof cost–effectively is beyond the state of the art
Testing
Alloy
Points
correct which sayspred correct {
there_are_such_things_as_points
} and I can ask Alloy to run this testrun correctand Alloy tells me that
The name "there_are_such_things_as_points" cannot be found. which is excellent news. I'm well on the way to using the familiar TDD cycle. Not compiling is failure and here is a failing test. I can make the test fail in a slightly more informative way by defining there_are_such_things_as_points like sopred there_are_such_things_as_points{
#Point > 0
}which says that the size of the set named Point (which is the set of all tuples conforming to the signature Point—it's a relational model, remember) is strictly greater than zero. Of course I haven't defined that signature yet so Alloy tells me that The name "Point" cannot be found. I define Point like sosig Point {}and now Alloy reports thatExecuting "Run correct" Sig this/Point scope <= 3 Sig this/Point in [[Point$0], [Point$1], [Point$2]] Solver=minisatprover(jni) Bitwidth=4 MaxSeq=4 SkolemDepth=2 Symmetry=20 18 vars. 3 primary vars. 23 clauses. 183ms. Instance found. Predicate is consistent. 41ms.Executing "Check check$1" Sig this/Point scope <= 3 Sig this/Point in [[Point$0], [Point$1], [Point$2]] Solver=minisatprover(jni) Bitwidth=4 MaxSeq=4 SkolemDepth=2 Symmetry=20 18 vars. 3 primary vars. 24 clauses. 11ms. Counterexample found. Assertion is invalid. 18ms.sig this/Point. In working with my model Alloy has made some instances of Point form which it has then constructed instances of the model. By default it chooses to make up to 3 instances of a signature. Here is a graphical (in both senses) representation of the instance of the model which Alloy built
Do you see this instance named in the array of three instances which Alloy reported it had created? Clearly the predicate is satisfied. Alloy will also produce a graph of the counterexample which it found—which is empty. (Well, strictly it's a message telling me that "every atom is hidden" in a "this page intentionally left blank" sort of way).Point. The problem domain can help us here, as it turns out that some of the points on the board have names. Here I state a fact, which is very much like a predicate, except that it is information for Alloy to use not a question for it to ask of the model.
pred tengen[p : Point]{}
fact tengen_exists {
one p : Point |
tengen[p]
}
tengen_exists like this: "it's true of exactly one instance, named p, of the signature Point that the predicate tengen is true of p". The predicate itself is parameterised on an instance of Point but does not depend upon that instance. Which seems as if it should smell.tengen[Point$0] which is (of course) true. If I ask Alloy to check the model it now reports No counterexample found. Assertion may be valid. 69ms. Note the "may be" there. Alloy can't be absolutely sure because it only instantiates a small number of tuples for each signature. This is a manifestation of the Small Instance Hypothesis (sometimes "small model" or "small scope") which claims that if your model is bogus then this will show up very quickly after looking at a small number of small examples—exhaustive enumeration of cases is not required.Points) to be invalid. I'll check in.Refinement
enum Name { Tengen }
sig Point {
name : lone Name
}
pred tengen[p : Point]{
p.name = Tengen
}
fact tengen_exists {
one p : Point |
tengen[p]
}Several new Alloy features are used here. Since Alloy 4 doesn't support string literals I use an atom (an instance of a signature with no further structure). The enum clause creates quite a complex structure behind the scenes but gets me the atom Tengen. The signature of Point is extended to have a field named name which will, in a navigation expression such as p.name resolve to an instance of signature Name, or to none, as shown by the cardinality marker lone. These navigation expressions look like dereferencing as found in OO languages, but are actually joins.
Directions
enum Name { Tengen }
sig Point {
name : lone Name,
neighbour : Direction -> lone Point
}
pred tengen[p : Point]{
p.name = Tengen
}which admits a counterexample which does not satisfy this predicate
fact tengen_exists {
one p : Point | tengen[p]
}
enum Direction {N}
pred tengen_has_a_neighbour_in_each_direction{
let tengen = {p : Point | p.name = Tengen} {
not tengen.neighbour[N] = none
}
} The counterexample looks like this
Direction -> lone Point which is pretty much the same as a typical "dictionary" and the let form and its binding of the name tengen to the value of a comprehension. The comprehension should be read as "the set of things p, which are instances of Point of which it is true that the value of p.name is equal to Tengen" Some sugar in Alloy means that we don't need to distinguish between a value and the set of size 1 who's sole member is that value.
This is an interesting state of the world, so I check in with a suitable caveat in the message.A YAGNI Moment
Point.direction. There are points on a Goban which do not have four neighbours, one in each direction. But I haven't mentioned any of them yet. There's a good chance that eventually points will need to have optional neighbours, but right now YAGNI.As the (so far, incomplete) predicate's name suggests, tengen really does have four neighbours. The cardinality should be
one. Making that change produces a model which no cannot be shown to be invalid. However, the example instance is still bogus. I know from TDD practice what to do: write a test that will fail until the problem is fixed. Here it ispred points_are_not_their_own_neighbour {
all p : Point |
not p in univ.(p.neighbour)
}The construction univ.r for any relation r evaluates to the range of the relation.As I hoped, this test fails. Although running the predicates can produce an instance in which the northerly neighbour of tengen is not tengen, checking can also still produce an invalidating counterexample in which it is. I must add a predicate to apply to all
Points forcing them not to be their own neighbourHere I use the function
sig Point {
name : lone Name,
neighbour : Direction -> one Point
}{
not this in ran[neighbour]
}
ran imported from the module util/relation to state the constraint on the range of neighbour. The conjunction of the predicates listed in curleys immediately after a sig are taken as facts true of all instances of that signature. The model is now consistent and not demonstrably invalid, but a glance at the example reveals that all is not well
This is interesting, so I check it in.
Complementary Directions
Once again, I need to strengthen the tests. If a point is the northern neighbour of tengen, then tengen is the southern neighbour of the that point. Directions on the board come in complementary pairs.pred directions_are_complementary{
N.complement = S
S.complement = N
} and now I have to de-sugar Direction in order to insert the complement relation. And now we see how enums workabstract sig Direction{
complement : one Direction
}{
symmetric[@complement]
}
one sig N extends Direction{}{complement = S}
one sig S extends Direction{}
In the fact appended to Direction I say that the relation complement (the @ means that I'm refering to the relation itself and not its value) is symmetric using the predicate util/relation/symmetric. Thus I do not have to specify that S's complement is N having once said the converse. The same pattern applies to E and W.The instance is now a spectacular mess.
I will check in anyway.Distinct Neighbours
pred neighbours_are_distinct{
all p : Point |
all disj d, d' : Direction |
p.neighbour[d] != p.neighbour[d']
}The nested quantification uses the disj modifier and should be read "for all distinct pairs of Direction, named d and d'..." run correct for 5 Point The resulting example is a rats nest of dodgy looking relations (it's in the repo as instance.dot if you want a look)pred neighbours_of_tengen_are_distinct{
let tengen = {p : Point | p.name = Tengen} |
all disj d, d' : Direction |
tengen.neighbour[d] != tengen.neighbour[d']
}
and with this I see a much less tangled, but still wrong, instance (instance1.dot). And the model is also demonstrably invalid. I'm going to make a significant change to the model. One I've been itching to do for some time. I check in before this.A Missing Abstraction?
InteriorPoint. I remove the predicates about distinct neighbours and introduce InteriorPointsig InteriorPoint extends Point{}and can then quite happily saypred interior_points_have_a_neighbour_in_each_direction{
all p : InteriorPoint {
not p.neighbour[N] = none
not p.neighbour[E] = none
not p.neighbour[S] = none
not p.neighbour[W] = none
}
}
and this gets me back to a consistent, not demonstrably invalid (although still wrong) model. I check in. sig InteriorPoint extends Point{}{
#ran[neighbour] = #Direction
}
and leave other kinds of point to look after themselves. If the range of the neighbour relation (which is a set) is the same size as the set of Directions, then there must be one neighbour per direction. This leaves me with the model in this state
sig Point {
neighbour : Direction -> lone Point
}{
not this in ran[neighbour]
all d : dom[neighbour] |
this = neighbour[d].@neighbour[d.complement]
}
sig InteriorPoint extends Point{}{
#ran[neighbour] = #Direction
}
fact tengen_exists {
#InteriorPoint = 1
}
abstract sig Direction{
complement : one Direction
}{
symmetric[@complement]
}
one sig N extends Direction{}{complement = S}
one sig S extends Direction{}
one sig E extends Direction{}{complement = W}
one sig W extends Direction{}
a little bit of tidying up and I check in. The model is consistent and cannot be shown to be invalid.
The example instance looks respectable too—so long as we focus on the interior point and don't worry about how its neighbours relate to one-another, which is clearly wrong. But there are no tests for that. This diagram has been cleaned up in omnigraffle to focus on the interior point but the .dot of the original is checked in. Maybe another time I'll sort out the regular points.Thoughts
Wow, that was hard work. Took a long time, too (although not so long as the timestamps make it look, I was also doing laundry and so forth during the elapsed). Does that make what is after all merely a rectangular array of points a "complex domain"? No. I'm out of practice with this kind of thing, and not fluent with the tool. Even so, I'm impressed by how good a fit the TDD cycle seems to be for this formal modelling tool. I even got into a bit of trouble towards the end but was rescued by recalling the TDD technique of making the tests dumb and repetitive—but concrete and clear. And how the same subtle trap of thinking too far ahead applies here too.Observations on Spotify
Bayesian Testing?
Introduction
Evidence
You may have read that absence of evidence is not evidence of absence. Of course, this is exactly wrong. I've just looked, and there is no evidence to be found that the room in which I am sitting (nor the room in which you are, I'll bet: look around you right now) contains an elephant. I consider this strong evidence that there is no elephant in the room. Not proof, and in some ways not the best reason for inferring that there is no elephant, but certainly evidence that there is none. This seems to be different from the form of bad logic that Sagan is actually criticising, in which the absence of evidence that there isn't an elephant in the room would be considered crackpot-style evidence that there was an elephant in the room.A Small Example of Confidence
Let's say that we wish to write some code to recognise if a stone played in a game of Go is in atari or not (this is my favourite example, for the moment). The problem is simple to state: a stone with two or more "liberties" is not in atari, a stone with one liberty is in atari. A stone can have 1 or more liberties. In a real game situation it can be some work to calculate how many liberties a stone has, but the condition for atari is that simple.if), but not greatly so. Laurent proposed a different question to ask from the one I was asking before—a better question, and he helped me find and understand a better answer.| One Liberty Means Atari | |
|---|---|
| liberties | atari? |
| 1 | true |
Adding More Test Cases
Suppose that I add another case that shows that when there are 2 liberties the code correctly determines that the stone is not in atari.| One Liberty Means Atari | |
|---|---|
| liberties | atari? |
| 1 | true |
| 2 | false |
| One Liberty Means Atari | |
|---|---|
| liberties | atari? |
| 1 | true |
| 2 | true |
| One Liberty Means Atari | |
|---|---|
| liberties | atari? |
| 1 | true |
| 2 | false |
| 3 | false |
| One Liberty Means Atari | |
|---|---|
| liberties | atari? |
| 1 | true |
| 2 | false |
| 3 | false |
| 4 | false |
That result seems at to contradict Dijkstra: exhaustive testing, in a case where we can do that, does show the absence of bugs. He probably knew that.
Next?
My brain is fizzing with all sorts of questions to ask about this approach: I talked here about retrofitted tests, can it help with TDD? Can this approach guide us in choosing good tests to write next? How can the structure of the domain and co-domain of the functions we test guide us to high confidence quickly? Or can't they? Can the current level of confidence be a guide to how much further investment we should make in testing?New article for BCW
Innovation Games
Sketches
6. MAKE MANY SKETCHESI think that this applies equally well to programming.
Join the best sketches to produce others and improve them until the result is satisfactory.To make sketches is a humble and unpretentious approach toward perfection.
—Fundamentals of Musical Composition, Ch XII
XP Day London 09: Programme
Scheduling by value?
value varies along an exponential scale while development costs vary along a linear scale. Therefore delivering the most valuable features trumps any consideration of whether or not the most valuable feature is cheap or easy to developwhihc, if true of your environment, might give pause for though. How this interacts with the desire to schedule so as to maximise throughput at the bottleneck is an open question, for me at least.
Service-Oriented Architecture
Observations on Estimation
While all of that is going on teams that want to use a numerical scale to estimate (rather than, say, "t-shirt" sizing) tend to choose a scale, a sequence of licit values from which estimates must be drawn. The various planning tools that demand a numerical field be filled in tend to force this issue.
I've noticed a tendency for "expert" level practitioners to want to use some clever non-linear scale, maybe Fibonacci numbers (1,2,3,5,8,13), maybe a geometric series (1,2,4,8,16) and they will have some sophisticated reason why this or that series is preferred. And I've noticed that a lot of teams aren't comfortable with this. They want to use a linear scale.
It seems to be traumatic enough that the estimates don't have units, or even dimensions. The idea that estimates are dimensionless but also structured can be a double cause of confusion.
Anecdote: a team had been estimating and planning and delivering consistently for a good long time. Their velocity was fairly constant, but drifted over time (fair enough). One day it turned out that their velocity happened to be numerically equal to the number of team members times the number of days to the next planning horizon. Someone noticed this and with a huge sigh of relief the team concluded that these mysterious "units" in which they estimated were actually man-days in disguise. Now they finally understood what they were estimating! And they promptly lost the ability to estimate: their next planning session was all over the place and it took some time for their planning activities to converge again. My inference was that it's actually quite important that estimates are dimensionless.
Anecdote: a User Experience expert at a client had been involved in some research whereby (as a side effect) members of the general public had to create a scale that made sense to them within which to rank the usability of features. These folks were presented with different generic objects and asked to give them a "size", and then to give a corresponding "size" to some other generic objects in order to create a scale that made sense to them, which would then be applied to the merit of the system features that were the actual target of the research. They created linear scales.
That surprised me at first, since I know that the physics of our sensory apparatus are generally non-linear, and memory is non-linear and so forth. But thinking about it some more I realised that our experience tends to seem to be linear, even if the underlying phenomena aren't.
Meanwhile, if one did want to use a particular scale for estimating the size of stories, why not use one of the series of prefered values? They are very well established in engineering and product design and offer interesting error-minimising properties. On the other hand, it might be a real struggle to get a team to decide if a story was a 1.6 or a 3.15
I don't have a grand narrative into wich to fit these observations, but here is another related anecdote about estimation.