Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Friday, April 26, 2013

Difficulties of the English legislator

I decided to take a look at what sort of reception was given to the 1841 Census in England – at least as displayed in the columns of The Times.

The situation had clearly changed completely from the late C18th controversies surrounding the very idea of a census, & the protests which greeted that of 1801. The idea of numbering the people once a decade had proved its worth.

Since I used just the word ‘census’ for my keyword search I brought up several reports where 1831 Census results were still being used to bolster argument – for example over the need for more schools, for public health initiatives & in relation to elections. There was intense interest in finding out by how much the count would show the population had grown over the decade – or if it had after all been reduced by emigration.

The one area which roused some difficulty was the date; originally set for mid-year (the night of 30 June/1 July), Census day was moved forward to 6 June after it had been pointed out that the end of the month would clash with the Quarter Sessions (the Courts which heard cases too serious to be dealt with by the local Justices of the Peace), which would mean an influx of visitors to every County Town, so artificially inflating the apparent size of their populations.

The system of having enumerators – “‘intelligent persons’ residing in the area” - to deliver & collect the forms to every household & then produce the first summaries by transcribing the information, within a week,  into notebooks worked well. They were paid 10 shillings for every 50 houses they dealt with (the usual number) plus 1 shilling for each 10 above that, up to a maximum of 80 households. They worked all the daylight hours on the day following the census, & some needed police escorts in ‘certain portions inhabited by the lower orders’.

By 26 July The Times was reporting some of the first results for London, with near final population counts available by the end of October.

Detailed analysis took longer of course. A geographically detailed Occupation Abstract was published in September 1844.

Confirmation of the continued fall in the numbers who derived their income solely from agriculture & the continued flow of population to the towns still left plenty of scope for political argument however in the period of intense debate about abolition of the Corn Laws, particularly the revelation that the majority were involved neither with agriculture nor cotton. There was surprise, & some puzzlement, that the population still managed to grow in size, by over 2 million since 1831, despite the drain of emigration. And marvelling at the fact that nearly 1 million women worked as household servants.

The Times leader pointed out that this analysis went someway to rectifying the great ‘want of authentic statistical information [which] is one of the greatest difficulties of the English legislator. In the absence of public documents he is driven, with more or less reluctance, to the suspicious estimates of theorists, the flagrant exaggerations of party, & the precarious guesses of absolute ignorance.’

Only some way however – there was still scope for arguing over accuracy & definitions – for example the total labour input to agriculture might be underestimated by allowing only one occupation pr person – a female servant might, after all, milk the cow.

Today of course mere information is not enough to meet the needs of the legislator who must have randomised controlled trials to resolve the flagrant exaggerations of ministers.

Link
The controversial introduction of the modern Census
Times archive
Related post
Numbering the dead

Wednesday, April 17, 2013

Big Victorian data

Some academic statisticians are feeling a bit worried about the future of their profession in an age of Big Data, Data Science, Data Analytics & Analysts

Big Data can mean working with billions of pieces of data - the initial challenge involves computer science rather than statistical formulae.

But is the challenge any more formidable than that which faced our statistical forebears?

In 1841 The Registrars General organised the first modern census of the population of the United Kingdom. Forms were delivered to every household to record personal details of over 25 million people. These were transcribed into notebooks by an army of local enumerators who collected & delivered the forms & helped those who could not read or write to fill them in.

I do not know if there were any kind of calculating machines available to help in the production of the rich variety of tables & analyses which were then typeset, printed & reported to Parliament. I suspect that most of the labour involved good old mental arithmetic undertaken by rooms full of Bob Cratchits. Even so, the logistical & organisational challenge must have been formidable.

A census – a complete enumeration of the population of interest – is not a sample survey whose results require the application of tests of significance & questions about how to make inferences to the population which has been measured. The challenge, before the vast modern increase in computer power & storage, was to specify which cross tabulations would be required out of the infinitely many which could be made.

In today’s world even a billion data points may be only a sample of the population of interest, which poses a new & different set of challenges.

Link
Simply Statistics: Data science only poses a threat if we don’t adapt
Significance Special Issue: Big Data
The UK 1841 Census
Related post
The first computer error











Wednesday, March 27, 2013

Hard data, hard times

Charles Dickens would not approve of Big Data; he had no time at all for Mr Gradgrind’s view that ‘You can only form the minds of reasoning animals upon Facts: nothing else will ever be of any service to them’, nor for his efforts to turn his son Tom into a scientist by having the boy trained to mathematical exactness.

Dickens had no respect at all for the output of Victorian collators & compilers of statistics in the form of Parliamentary Blue Books.

Whatever they could prove (which is usually anything you like), they proved there, in an army constantly strengthening by the arrival of new recruits … the most complicated social questions were cast up, got into exact totals, and finally settled - if those concerned could only have been brought to know it. As if an astronomical observatory should be made without any windows, and the astronomer within should arrange the starry universe solely by pen, ink, and paper, so Mr. Gradgrind, in his Observatory (and there are many like it), had no need to cast an eye upon the teeming myriads of human beings around him, but could settle all their destinies on a slate, and wipe out all their tears with one dirty little bit of sponge.
Heaven knows what Dickens would think of the idea of settling issues of social policy through the mechanism of the randomised controlled trial
Well, actually he gave us a pretty good clue, in the words of young Tom Gradgrind, failed mathematician:
'I wish I could collect all the Facts we hear so much about,' said Tom, spitefully setting his teeth, 'and all the Figures, and all the people who found them out: and I wish I could put a thousand barrels of gunpowder under them, and blow them all up together! '

Those Blue Books are now regarded as invaluable social documents by many, but there will always be those who  place more value on the story than on statistics.

Link
Project Gutenberg: Hard Times by Charles Dickens

Saturday, March 16, 2013

Arbitrarily at random


A recent Times crossword (#25410) used arbitrarily as the definitional part of a clue; the answer required was at random.

I was inclined to quibble: haphazard, unarranged, accidental, maybe. But arbitrary suggests just personal whim.

The OED seemed to provide some support. Arbitrary does not appear in any of the definitional entries for random, & random does not appear under the definitional parts of either arbitrarily or arbitrary.

Interestingly, however, according to the OED, random play is ‘a facility on a digital music player for playing tracks in an arbitrary order’. One wonders just exactly how the algorithm works for that choice.

Even more interestingly, the word random itself comes originally from Norman French & means speed, haste, impetuousness, violence, or to run fast, gallop. Hence, presumably, the lack of time for thinking & deliberating on your random choices.

Related posts
Getting my knuckles in a twist
Humpty Dumpty on statistical sampling

Saturday, March 02, 2013

Ministry of Ag & Fish


Nate Silver does not like Sir Ronald Aylmer Fisher. Not his style of dress (ill-kempt & dishevelled), nor his combative, argumentative style; Silver believes there must be a uniquely British mould for producing such intellectuals – after all we also gave Christopher Hitchens to America.

But most of all, Silver detests Fisher’s intellectual legacy, blaming him for the widespread adoption of the purely frequentist approach to statistical inference & the neglect of the theorem of Thomas Bayes.

As if more were needed, the most devastating charge against the man is that he continued to dispute the relationship between smoking & lung cancer, maintaining that researchers who claimed to have established a causal link were confusing correlation with causation.

Although Silver writes passionately about his own belief in the Bayesian approach it is not at all clear to me from his book that he uses formal Bayesian methods, with specified ‘priors’ in his own predictive modelling. Rather he advocates the use of judgement & deep knowledge of the process(es) underlying the forecast to adjust, where necessary, the predictions to take account of uncertainties which can never be reliably & completely captured in a mathematical black box. And in his concluding chapter he states baldly that ‘even common sense can serve as a Bayesian prior’. What is unacceptable, he says, is to believe that mathematical forecasts have a purity entirely free of the forecasters own subjectivity & biases, which do not need to be acknowledged.

I hold no brief for Fisher & am very much in sympathy with the Silver approach, but I am taken aback by the vitriol of his attack upon Fisher, even though I too despair about the cult of significance which he could be said to have inspired. Fisher was more subtle than that &, like Marx, can be said to suffer from the quality of his disciples. It is also important to remember that Fisher –born 1890 – was developing methods which could be used in an age when a computer was a human calculator.

I smiled at the bit where Silver – born 1978 - writes about how a standard SIR model in epidemiology takes only seconds to run on a laptop. Provided of course that somebody else has already done the hard & expensive work of collecting & collating all the data & making it available in electronic form at a click of a button to any one who cares to use it.

I cannot remember ever seeing any of Fisher’s contributions to the debate on smoking & lung cancer, but I was privileged to hear Sir Richard Doll speaking on the subject several times. And he used to point out that, when he started to investigate lung cancer nobody expected tobacco to be implicated.  Some aspect of general air pollution – such as  the Great London Smog of 1952, credited with causing thousands of deaths – seemed much more likely to provide an explanation. They were also of course well aware of the problems of disentangling causation from correlation & were applying the Bradford Hill criteria, with a battery of approaches.

I was once present when someone asked Sir Richard at what point he personally had become convinced that smoking was the culprit. He replied he thought that it was probably when they got the results of a prospective study of patients referred to UCH for ?lung cancer? In every case where the diagnosis was confirmed the victim was a smoker; non-smokers were all found to be suffering from a different problem.

As far as I am concerned – the debate was settled once & for all in March 1962 when it was announced on the BBC 6 o’clock News that it was now ‘beyond doubt’ that smoking causes lung cancer.

Nevertheless even in 1966 an undergraduate dissertation on ‘The Relationship between smoking & health’ would not have got good marks had it failed to mention some of the outstanding doubts or queries.

This was a dissertation for a [service] course in Applied Statistics for those doing degrees in economics & so was, in effect, a discussion of statistical, rather than medical issues. Sample selection (the largest UK study looked only at male doctors), cross-country comparisons, looking at rates of smoking & rates of lung cancer, etc. (As I remember it Australia & Switzerland did not fit the pattern which otherwise showed a positive correlation between rates of smoking & rates of lung cancer; both had rates of the disease which were much lower  than expected.). I also remember – bizarrely as it now seems – concerns about whether people who stopped smoking might gain weight & therefore put themselves at increased risk of a heart attack.

And – most important of all – what could explain the fact that the majority of smokers did not get lung cancer – we were not allowed to consider genetic explanations in those days, so factors such as type of cigarette, the method of inhalation, the precise make up of the tobacco etc were candidates.

Although I don’t think I knew anybody who seriously doubted that the link had been proved, there were many who thought that what was really needed was an answer to that question. The real scientists & medics must grasp the baton handed to them by the statisticians.

I do not know what Fisher’s own reaction had been to the publication of the report by the Royal College of Physicians. He did not have very much time to continue the argument he died unexpectedly on 29 July 1962 in Adelaide from an embolism after a successful operation for bowel cancer.

Links
Smoking and health (1962)
Related post
Pretty girl in crimson rose


Thursday, February 28, 2013

Sorting out average pay

The recent report from the LSE Growth Commission recommends that median household income should take a regular place alongside GDP as a measure of how the economy is doing.

That brought back memories of the 1970s, when governments still often assessed economic effects of the Budget by quoting the cost (or benefit) to the average family with two children & a father on average earnings. Average was always just understood to be the arithmetic mean & as such a fair representation of how much ‘most’ ordinary families had to live on.

As the great inflation following the oil price shock started to drag more & more workers into income tax & the Government tried to calm union demands by referring to the ‘social wage’, one Labour MP – I think it must have been Audrey Wise – started to table a series of Parliamentary Questions asking for comparisons of median with average earnings. Some of us were quite impressed – in an age when knowledge of even basic statistics was spread so narrowly – that someone was alert to the existence of different kinds of average. Those of us who already knew that earnings/income do not generally follow a symmetric, bell-shaped curve (or normal distribution) but one which is skew, with a long tail towards the higher end of the income scale were even more impressed. On a bell-shaped curve [arithmetic] mean & median have the same value, but (right to left, reading up the hill) for a positively skew unimodal distribution the mean always exceeds the median. My memory tells me that the median was, in those day, only about 2/3 of the mean for earnings of men in full-time work.

Such lack of education on this point continued for well over a decade – I remember trying to explain to someone how it was that Mrs Thatcher was not (necessarily) telling lies when she talked about the new poll tax in relation to ‘average earnings’ of around £10,000 a year – ‘Nobody round here earns anything like that’.

In 2013 we seem to have gone the other way – I have been finding it hard to lay my hands on average (mean) earning figures for the UK – medians are everywhere now.

One reason why medians did not get so much attention in those days may have been that they could be difficult (or at least very expensive) to calculate from large surveys. Records were held in fixed format in fixed order on magnetic tape which could be read in only one direction. Sorting could take forever. Astonishing how we now take it for granted that sorting (of a sort) can now be done easily at the click of a button.

Links
LSE Growth Commission
Cabinet Papers: The Miners Strike and the Social Contract
[PDF] Introduction to sorting
YouTube: [Nerdy] Sorting Out Sorting
Related post
Assorted memories

Monday, February 11, 2013

Probably a percentage

In three papers this last week it was interesting to see words such as probability or risk used to describe what I would probably have called, simply, percentages. I might have added terms such as growth, change, share, and concentration, distribution or simply more or less, higher or lower to aid description, analysis or explanation.

This probably reflects my training in economics ‘analytical & descriptive’ in the days before the subject became dominated by mathematical models, followed later by work in data collection via censuses & (mostly social) surveys; this latter day shift towards thinking in terms of probabilities may reflect a more widespread acceptance of the idea that we live in a probabilistic world, or just today’s emphasis on model-building & parameter estimation.

It would be interesting to trace the history of when percentage breakdowns (as in the old joke Population broken down by age & sex) became the norm in the presentation of statistical tables. The early authors of papers in the Journal of the Royal Statistical Society use simple proportions with varying denominators, comparing 5 out of 9 with 7 out of 12, for example. Presumably Victorians were better at knowing which was higher.

The British Hedgehog Survey illustrates population decline with an index (setting the base year to 100), which is interpreted as ‘a measure of the probability of detecting a hedgehog in a particular year relative to the probability in the first year of the survey.’ This set me pondering whether there might be anything to be gained by thinking of, for example, the Consumer Price Index, as a (weighted) probability that any product you put in your trolley at the next visit to the supermarket will have increased in price since last month, rather than as a measure of the ‘thing’ we call inflation.

An engagingly written report Hilary: the most poisoned baby name in US history looks at the risk that a name might suddenly fall out of favour with new parents making that all-important decision about what to call their baby. The author talks of ‘a measurement called the relative risk, where “risk” refers to the proportion of babies given a certain name’ and, even more engagingly, explains how to calculate percentage rates of growth and decline.

The third example came from a magazine article about knife crime in London, which quoted a study of the risk that youngsters living in south London might suffer from various kinds of deprivation such as living in a lone parent family without a father, or poor educational attainment. Here the concept of area-based risk comes closer to implying a cause & effect relationship – the area causes the deprivation.

It is not so long ago that the same sort of statements could have been made about Notting Hill. I wonder if the millionaires & Cabinet Ministers who live there now have any such fears for the future of their own children?

Links
PDF]The state of Britain’s hedgehogs 2011
Hilary: the most poisoned baby name in US history
Related posts
Social mobility
The verb to be
Onwards & upwards
Making things add up

Saturday, February 09, 2013

Public library lending


The annual ‘best seller’ list for UK public libraries has just been published.

Much was made of the enduring popularity of the novels of Danielle Steel, whose books have been featured in the lists for each of the last 30 years. Her books were borrowed almost a million times last year alone.

Jack Malvern, writing in The Times, interpreted this as the ‘British reading public’ showing what they think of literary critics & cultural commentators who mock her ‘formulaic stories’.

Well maybe.

But that figure needs to be seen in the context of the total number of loans made by UK public libraries last year.

The novels of Danielle Steel accounted for  0.3% of a total of more than one quarter of a billion loans (287,505,000 to be more precise – rather more than the 209,000,000 books bought by consumers in 2011 according to figures from the Publishers Association).

The best seller lists are compiled from figures which are used to make Public Lending Right (PLR) payments to authors. From which we can calculate that, with a total of £6.4m at a rate of 6.2p per loan, only some 100 million (roughly one-third of all library loans) attract the payment. Part of the gap is explained by the fact that there is a cap of £6,600 on the payment to any one author (ie they are not paid for more than about 1 million loans of their books in any one year), and the rest is made up of books by authors who do not qualify for PLR.

Links
What is PLR?
CIPFA national library survey

Friday, February 08, 2013

Tickled my fancy

Recent links I liked

The great GIF debate explains, among other things, why Jif (the sink cleaner) had to change its name – to avoid confusion with a brand of peanut butter

Issues with reproducibility at scale on Coursera reproducing someone else’s code or statistical analysis is not as straightforward as you think – or hope it is, or should be.

Staying Ahead of the Tax Man Everybody does it if they can.

Word String frequency distributions - Always check what the numbers mean before you analyse them

Investigating paediatric pain: Q & A with Professor Maria Fitzgerald and
Pain: Anaesthesia and Self-mutilation in the Nineteenth Century One disturbing, one just discombobulating, exploration of past medical views about pain

Wednesday, December 12, 2012

Some journalists do count


The 1971 Census in England & Wales was the first to be processed entirely by computer – though the forms had been delivered to & collected from every household in the traditional way by an enumerator, who first had the job of compiling a complete list of residential addresses in the area allocated to them.

The results were somewhat delayed.

Not solely because of computer problems – compared with the disasters to which we have grown all-too accustomed in recent years things went well - but delays were compounded by the publication process which stuck, for the most part, to doing things the old-fashioned way with bound volumes of printed tables produced according to a schedule fixed in advance. These volumes had then to be ‘Laid’ before the House of Commons before anyone outside the statistical office could see them, & this requirement led to the frustration that sometimes the figures were off the computer but could not be used publicly in any way until they had been laid.

In another measure of just how far we have come since those days, the first published results came out county by county, reflecting the way that the data was first processed one county at a time, with national analyses following on behind.

As publication day for the first county finally loomed I, for one, looked forward to seeing what, if anything, the national media might make of it, especially as it would be my own home county of Derbyshire which was in the spotlight. Most likely, I thought, it would be of no interest to the wider world.

So I was excited to see a whole column in The Guardian, carrying the byline of one of their most senior & respected reporters.

I was genuinely shocked when I realised that this column simply reproduced (albeit with a top & tail) the Press Notice which had been written by the (then) OPCS & sat gathering dust until it could finally be released. I had naively expected that a journalist would look through the volume to find his own story.

Forty years on & how things have changed. All newspapers produced their own charts of the first detailed national figures released this week, with even more details & analysis available on their web sites.

Data journalism is the new big thing – there is even a handbook for it which can be downloaded for free from the web

Links

Related posts


Friday, March 23, 2012

Lessons in probability

An interesting take on calculating the Bayesian probability that the sun will rise tomorrow from Allen Downey on Probably Overthinking It - The sun will probably come out tomorrow


And a variation on (corollary of?) Simpson’s paradox operating on equal pay, from Nigel Hawkes on Straight Statistics - Squinting hard at the gender pay gap


Thursday, March 08, 2012

Laws of statistics

Statistical information is not real information. A statistical description – this includes models – is, like a painting, only a representation of reality, although, as with a portrait, some have more ‘likeness’ than others.

Statistics apply to populations, not individuals.

Post-selection randomisation is not the same as selecting a random sample from the population you wish to study.

Increasing the size of a biassed sample reduces sampling error, but not bias. So the ratio of bias to sampling error increases.

If P is set at 5% then expect 1 in 20 of your tests to give ‘significant’ results even when H0 is true

An insignificant result is God’s way of telling you your sample is too small

Saturday, February 25, 2012

Bins & buckets

Binning data seems a very wasteful thing to do. Even if you have finished with it, someone else may find it useful.

It was when I was recently reading a paper written by a physicist about statistical distributions that I first came across the concept of binning data. I had to look at the figures before I understood what he meant, & then more minutes to try & think what word I would have used. Probably grouping, or maybe just cross tabulation which, in social statistics usually involves grouping - for instance population in 5-year age bands.

I thought binning must derive from the idea of sorting data into different bins according to the value of the variable – much as, in the old days, postal workers used to sort outgoing mail by tossing it into the correct bin for the postal town of destination.

Just the other day it occurred to me that bin may in fact derive from binary, which corresponds to the truth table method of programming cross tabulations.

And then I came across this in an article about statistical analysis of language & economics on Language Log:

What this does, in effect, is drop families around the world into one of 1.4 billion buckets, where two families fall into the same bucket if and only if they are identical in country of birth and residence, age, sex, income, family structure, number of children, and religion, where the religions of the world are broken up into 74 types.
Keith Chen


So perhaps I was right first time.

Saturday, January 14, 2012

Open data

On New Years Eve Tim Berners-Lee had an opinion piece in The Times about open data:
“All this data has been paid for by taxpayers. So the … mission will be to make sure that we can all make the best use of it”
The argument that the data is already paid for is interesting & used to form the basis of the pricing policy for Government publications; in the days when printed paper was the only option. marginal cost pricing meant covering the costs of printing & distribution only, content was free at that point to the user. As an undergraduates we could be expected to furnish ourselves with a copy of some relevant Government White Paper or statistical digest which, from memory, generally used to cost less than 2/-, cheap even in 1960s money.

As far as Government statistics were concerned however this came to an end with the publication of the Rayner review of 1981.
“There is no more reason for the government to act as universal provider in the statistical field than in any other”
Cmnd 8236: THE GOVERNMENTS STATISTICAL SERVICES.
And that wasn’t all: rather than provide data at marginal cost we were to start charging what the market would bear. Which of course posed a problem when the market may, in part, consist of students or members of the public keen to participate in debate as informed citizens of a democracy, & in other parts of businesses to whom information offers great potential for profit & who can therefore afford to price small none commercial users out of the market.

Wednesday, January 04, 2012

Doctored


An article in the Christmas double edition of The Economist - The disposable academic: Why doing a PhD is often a waste of time
– has been attracting a lot of attention.

When I was an undergraduate it was not very common (outside the natural sciences) for anyone to do a PhD as a full time student. Those destined for an academic career could still graduate in July & start work as an assistant lecturer in October (with the prospect of acquiring tenure in two to three years time), something of which I remind myself when we hear complaints of modern students being taught by mere postgrads, not proper professors.

Of course research & publication were required, but formal theses, submitted for examination, were quite rare.

We knew they did things differently in America.

One story which was current in my day illustrates this.

Alan Stuart was assisting Maurice Kendall with the new, revised edition of The Advanced Theory of Statistics, which even today, (having gone through even more revisions & expansions & acquired even more authors) is the bible of the subject.

Kendall was then working at an American university & wanted Stuart to join him. In those more expansive days – no problem, we can find him post.

Except that when he arrived & was discovered to have no PhD – whoops, sorry, our rules don’t allow mere bachelors to hold proper jobs in the academy.

The laid back Brits solved that one easily enough.

The Senate of the University of London simply awarded Stuart a DSc, which was clearly merited by the quality of his work.

I have been unable to trace any mention of this, quite probably apocryphal, or at least embroidered, story on the public web via Google, nor any reference to Kendall and/or Stuart spending any time at an American university during the years of the preparation of the new edition the late 1950s, though it is possible that the obituary in the Journal of the Royal Statistical Society Vol. 48, No. 2, 1999 may provide some kind of confirmation.

Saturday, October 29, 2011

Aliens desirable & undesirable

Much was made this week about the news that, out of every seven people in prison for offences connected with the August riots, one was a foreign national.

I wonder what those same commentators make of the fact that, in many of the poshest areas of London, a very high proportion of residents (a lot more, I would guess, than 1 in seven) are also foreign nationals, born ‘as far away as Cuba, Samoa & Vietnam’, not to mention China, Russia & the Middle East.

Why are commentators much less concerned about the foreignness of these people & the effect they have on our way of life & property prices?

Even David Aaronovitch concluede that the lastest figures just go to show that the rioters were just the usual suspects.

Too right they were - three-quarters of all those who appeared in court had a previous conviction or caution. For adults the figure was 80% and for juveniles it was 62%.

But this cannot be taken to represent rioters as a whole.

Given that you were a rioter, the chance of getting identified from film & cctv footage & picked up by the police, was much greater for those ‘already known’. Those who were ‘not known’ had a much greater chance of melting away.

Thursday, June 02, 2011

No initial significance

It began with a conversation about the wisdom or otherwise of going for an interest only mortgage – something which concerns my generation in relation to their children’s long term financial security. Reminiscences of how things used to be, when the main worry might be to ensure the mortgage was paid off if daddy died before the family were grown.

Another one of those changes brought about by the astonishing increase in longevity. Although some of us actually have a parent or parents who are still alive, have already received a telegram from the Queen, in our childhood it didn’t seem rare or unusual to think that a man might die in his forties.

We then turned to thinking about people we had known – of our own generation - who nevertheless died untimely in this way.

I had just said that I could, without thinking too much about it, immediately call to mind two former colleagues – one man, one woman – who had died (non-accidentally) aged 45 & 43 respectively, when it suddenly struck me that both had the initials KM.

Now I can do the calculations; if you take the line that the first death sets the bar, this is just proportional to the probability that someone born in 1940s Britain had the initials KM, assuming that the probability of death in the 5th decade of life is independent of name & ignoring the complications of name change on marriage etc, & any bias in my choice of friends or colleagues.

Not long odds.

So why do I find it difficult to shake off this funny feeling about it.

Related post
Time of birth

Saturday, April 16, 2011

Irish Census

It was census Day last Sunday in Ireland. Today I heard an advert on RTE Radio 1 reminding people to have their form ready for collection – there was no mention of any online alternative.

But there was a thank you in advance to everybody, ‘For Making Your Mark’

So much better than threats of a fine

Link
Irish Census

Related post
Census

Thursday, March 31, 2011

Census

There is quite a mood of pessimism around about the likelihood of a good response to the Census, which cannot be helped by the formidable-looking form which many must find daunting.

Then there is the clearly recognised problem of whether it is even possible to attempt to number the people on a single day in a modern society where nobody ever stands still, not just geographically but socially & in their personal relationships – a problem which may have contributed to the alleged undercounting of the population of New York in last year’s US census.

I have not been terribly impressed by the official poster campaign – fill in your Census form & help plan local services – that’s a policy or political argument for prioritising the expenditure. You don’t need statistics to make decisions about how to spend (or cut) money on services – I could make them right here, right now - & you would have a hard time proving that decisions based on Census statistics are better than my kind. Does anyone really believe that any mistakes made in the late 1970s were caused by the late cancellation of the 1976 census?

It strikes me that one reason why a recent survey found that a large percentage of the population does not even realise that a census is necessary to count how many people live in this country is that they no longer know the story of ‘the first census’, the one that took Joseph & Mary to Bethlehem. As children the link was made explicit for us, so that the modern census linked us to the birth of Jesus.

But then perhaps it’s just as well that today’s children don’t know the King James’s version: And it came to pass in those days, that there went out a decree from Caesar Augustus that all the world should be taxed*.

In retrospect it might have been better to present the Census as a real Big Society project, rather than an expensive way of collecting data that companies like Experian (a credit reference agency) supposedly already have.

The census is a collaborative effort to paint a statistical picture, a snapshot, of how we live now. In return for our cooperation we get a guarantee of 100 years of confidentiality for our individual personal details, which will then however be released to give those who come after us invaluable information about their forebears, the history of the house they live in, their neighbourhood, & even the domestic details of the lives of our celebrities.

Our census this year was done & dusted on line. It was a pleasant experience – ONS have shown that they can produce a really clear & well designed web site when they put their minds to it.

The only drawback was that I couldn’t get away with not answering the ethnic question, no box for ‘declined to answer’, so I had to use my ingenuity.

The wording of the question is interesting: Tick one box to best describe your ethnic group or background. Best in what sense? To whom or for what purpose, or in whose opinion? Can’t remember the wording last time but it is certainly different from 1991, which asked for the respondent’s own opinion. That of course is unsatisfactory, if we are hoping to get an insight into the extent of disadvantage, since discrimination depends on the observer’s assessment of ethnicity – which may be inaccurate, even bizarre.

The overall balance of the questions as usual reflects the political preoccupations of the last few years – we will have a comprehensive picture of ‘identity’ – if the response is good enough. Race, nationality, language, partnership status.

For me the most interesting change is the attempt to get at the detailed relationships within the household, recording the link of each person to all the others, instead of just to Person 1. So statisticians should not have to scratch their heads (which are wrapped in wet towels) about problems such as:

Person 1: Mr Smith, 50
Person2: Mrs Smith, 47, his wife
Person 3: Miss Smith, 18, his daughter
Person 4: Master Smith, 2, his grandson

Is Miss Smith the mother of Master Smith?

But how sad to read the detailed instructions in the nortes about how to count children who divide their time between homes.

*PS Word’s grammar checker suggests an amendment to the King James’s version. It would prefer that the entire world should be taxed.


Thursday, March 17, 2011

Heteroscedasticity

I have been meaning for some time to note that the Oxford English Dictionary does in fact help with the origin of the word heteroscedasticity – if you spell it correctly, which I was not doing when I first went looking for it; the variant heteroskedacity has not yet been recognised.

By 1901 the great Karl Pearson was devoting himself to full time research and teaching in the new mathematical field of statistics, with a grant from the Worshipful Company of Drapers which enabled him to establish the biometric laboratory at University College, London. Projects ranged from advances in statistical methodology to the analysis of data on heredity and physical anthropology as well as work on non-biological topics such as astronomy and dam construction which demonstrated the wide applicability of his new analytical techniques.

The first recorded mentions in print of heteroscedasticity & homoscedasticity, words coined by Pearson, came in 1905, in the pleasingly named Drapers’ Company Research Memoirs (Biometric Series):

If all arrays are equally scattered about their means, I shall speak of the system as a homoscedastic system, otherwise it is a heteroscedastic system.

Scedastic comes from the Greek σκεδαστ-ός,’ capable of being scattered’ & so heteroscedastic means 'of unequal scatter or variation; having different variances.' The word has only ever been applied to statistics.

The use of the word scatter left me wondering what measure was being used – when did standard deviation or variance come into use.

According to the OED the word deviation was first used by W. W. Greener in 1858 in his book on one of the most popular sporting pastimes of the day, Scientific Gunnery:
The mean deviation on the target from the centre of the group of 10 hits being only •85 of a foot at 500 yards' range.

but Pearson gave the first rigorous definition in 1894 in Philosophical Transactions of the Royal Society:
Then σ will be termed its standard-deviation (error of mean square).

It appears to have been left to R. A. Fisher in 1918, in Transactions of the Royal Society of Edinburgh, to introduce variance as a formal technical term in statistics:
It is desirable in analysing the causes of variability to deal with the square of the standard deviation as the measure of variability. We shall term this quantity the Variance.