Sunday, 31 August 2014

Identifying hoverflies from photographs



Identifying hoverflies from photographs is not straightforward. Photographs depict a single plane and do not allow rotation to look at particular characters. By comparison, preserved specimens can be viewed from many angles under the microscope. In addition, most keys have been developed from preserved specimens, which can differ markedly from live animals. The art of identification from photographs is therefore in its infancy. As photographic techniques improve, identification techniques may also improve; but there will always be some hoverflies that cannot be identified at all from photographs. That said, it is realistic to assume that around 50% of the British fauna can be identified from at least some photographs. There are a number of obvious ways in which photographers can be more assured of a positive identification:



  • The higher the resolution of the photo, the more chance of actually seeing key characters. For example, really nice sharp and well-filled frames can expose hairs on eyes, leg hairs and occasionally the pilosity of the arista.

  • Views from several angles top, front face and side view often combine to provide enough information to give a positive ID.

  • It is worth developing a knowledge of the family so that you have a rough understanding of the genus you are photographing. Each genus depends upon a slightly different range of characters and once you have a feel for the genus it should be easier to make sure that key characters are depicted.

There are some genera that cause particular problems - the most frequently illustrated genera that don't get identified are within Cheilosia, Eristalis, Platycheirus and Syrphus. That is not to say that they are the most difficult to ID in other circumstances, but they are the most frequently depicted genera from awkward angles. In addition, there are some tribes and families that are most unlikely to get identified because they rely on characters that are difficult to show in photographs. These include the Platycheirus where pits on the underside of male tarsi can never be seen in live specimens, and Eumerus, Pipizella and Sphaerophoria where it is not possible to examine the male genitalia. The tribe Pipizini is altogether difficult, even under the microscope and is unlikely ever to be readily identified from photographs; so too are many Cheilosia.

But, if one ignores the problems (a good idea) IDs can probably be given in 60% of cases to around 150 species. The others fall into the too difficult group or will only be ID's from an exceptional photo.

The following is a tabulation of the commonest ID problems that I encounter:

My ID (not necessarily right!)
ID posted
Eristalis intricarius
Criorhina berberina
Volucella bombylans
Eristalis pertinax
Eristalis tenax
Eristalis sp - various often not possible to go further
Eristalis tenax
Eristalis pertinax
Eristalis sp - various often not possible to go further
Eristalis sp.
All sorts of views, often at angles that show few characters or are well out of focus.
Eristalis rupium (quite regularly on iRecord)
Syrphus sp.
               
Epistrophe diaphana
Eupeodes latifasciatus
Megasyrphus annulipes - several on iRecord
Syrphus ribesii - the chosen name for about 90% of posts, suggesting that little attention is paid to text in the main keys or that Chinnery is being used.
Parasyrphus sp.
Eupeodes sp.
Xanthogramma pedissequum (agg)
Eupeodes corollae
Eupeodes luniger
Eupeodes latifasciatus
Parasyrphus punctulatus
Eupeodes luniger
Eupeodes corollae
Unidentifiable Eupeodes
Eupeodes luniger
Eupeodes corollae
Eupeodes latifasciatus
Leucozona lucorum
Volucella pellucens
Cheilosia illiustrata
Merodon equestris
Volucella bombylans
Platycheirus albimanus
Platycheirus scutatus
Platycheirus scutatus (agg)
Platycheirus albimanus - a problem I think resulting from the WILDGuide that I hope we will rectify in edition 2.
Scaeva pyrastri
Eupeodes luniger
Volucella pellucens
Leucozona lucorum
Xanthogramma pedissequum (agg)
Eupeodes nitens - a Chinnery mistake

Sunday, 24 August 2014

Records from photographs - a conundrum


At the start of this week I hit a brick wall in terms of the effort I make to extract data from websites. It is a pretty huge task these days, but when I started six years ago it was in its infancy. The growth in photographic recording has been massive and these days I am unable to keep on top of it without working a minimum of 50 hours a week. That is unsustainable as I find it getting in the way of my ability to earn a living! So I posted on various forums that I was going to cut my commitment at Christmas and would restrict my involvement to one or two forums.

I was looking to see what expressions of interest there might be to help out. The result has been very helpful, with an excellent software engineer volunteering to build a data extraction bot and website (he has one up and beta-testing already!). Two other people have offered help. On the whole, comments have been very positive but I did get one response that raised all the questions that are raised by taxonomic specialists – what is the value to be had from ad-hoc photographic records?

This is a really important question and one that deserves careful analysis because in my view the national datasets are increasingly skewed towards such data. Certainly that is the case for hoverflies and I suspect that a similar situation will obtain elsewhere. Does it matter? And, if it does not matter, what are the benefits of growing a bigger network of recorders who perhaps only record part of the fauna?

At one time I might have held similar views to those expressed above, but I have given a lot of thought to the issue and have concluded that on the whole the benefits vastly outweigh the drawbacks. My reasoning is as follows:

One can either take a highly insular approach to recording and confine recording schemes to the outputs of a very small number of recorders who cover all taxa within a particular family. Alternatively, one can absorb all records and recognise that the dataset will be disproportionately skewed towards those species that people see and can identify without resorting to taking a specimen and undergoing microscopic examination. The two approaches yield very different data profiles, and in the past the outputs of key recorders would have dominated the dataset (for about 30 years the dataset was dominated by just 20 very active recorders. Many of those recorders are no longer very active and the datset is now growing from a new cohort of recorders, rather fewer of whom cover all taxa. Thus, the HRS dataset now fits much closer to the dataset emerging from photos. All the same, provided one has a clear picture of who records in particular ways one can split the data according to technique and analyse it accordingly. So there is no real problem from a data management perspective.

There then comes the issue of rarity or difficulty of ID. Whilst the occasional record of a 'rarity' might be of some interest, it is of limited value when wanting to analyse trends – you need an awful lot of records to do much trend analysis, and by its very nature rarity precludes such analysis. In actual fact one does find from some photographs that there are more of certain species than we might think – e.g. Palloptera muleibris turns up far more frequently as photos than it does in my net (I think I've seen it twice in 30 years!). The data for many hovers and larger brachycera have contributed to the various species status reviews. So, if we judge datasets on rarity then maybe they are not covering all taxa, but in actual fact photographers do see species that the specialists rarely see – for example I reckon that there are more photographic records of Actophila superbiens this year than will come from specialists.

I would then suggest that 'common' species are often the bellweather of changes in the wider countryside, so big datasets of species that people can identify may actually tell us quite a lot about the natural world. We can do this with a variety of hoverflies from photos – changes in emergence periods and in distribution. The data are too limited yet to look at trends but they are improving.

Finally, I think we must look at what one is trying to do when engaging with photographic recorders. We have to be realistic that this is the biggest cohort of natural historians and is increasing in influence. We either engage and hope to show how there remains a need for sound recording by collection, or we shrink into a box and fight people off when attacked. I favour the former and that is why I put effort in. What is more, if the biological recording community is to remain active and relevant, the photographic community is a very big constituency so we need to engage and to show what can and cannot be done with data accumulated this way.

So, am I wasting my time? Well if some people think that is the case then they don't have to get involved. But, we rely on a very small band of people to make the detailed datasets and those alone will not provide some of the data that are important. If by outreach we pick up the occasional person that gets more deeply involved, or converts to using a microscope (there have been a few), then there is a future for sound taxonomic recording. If we fail to do that outreach and to show value to what people are doing, not only will recording diminish as the current generation pops its cloggs, capacity to generate new competency will also diminish. If I was a politician reading some comments I would be thinking – why bother with this lot – they are not inclusive and are negative. If I got too positive a message I might develop too many expectations.

My view and approach is to look for a level of input that demonstrates both the value and the limitations, but the main limitation is the lack of specialist capacity. I have previously written about the need to develop more expert capacity, and I believe that developing the recording effort via engagement with photographers is a valid way of doing so. It may not be 'ideal' but then if you wait for the ideal situation it is unlikely to happen.

It is also worth bearing in mind that the sort of engagement I have made means that large numbers of people have developed an interest in diptera at some level. Some will buy the books - e.g. the revised Larger Brachycera book. We need to drum up interest in order to sell enough books to make them economically viable. If we don't then such books will not get published. Increasing interest helps to sell the Wildguide and that in turn generates an income to produce guides such as the hoped for Scathophagid book (we have donated the proceeds to Dipterists Forum). Likewise, that interest may help to generate the case for a Diptera Wildguide - and for that my database may be essential to source relevant photos. So, it is not a simple question of limited records, there is a strategic case too.

There is a genuine need for debate about the value of different recording techniques, but the most important issue is to think about what positive benefits can be accrued. If we don't make an effort to extract data and to engage to encourage, then we are missing a time-limited opportunity, as the people who can provide the taxonomic expertise are aging and we need to grow a new constituency of specialists to provide the detailed taxonomic advice.

Sunday, 3 August 2014

Is something happening with Rhingia campestris this year?


Rhingia campestris is a common, readily identifiable species that is reported by specialists and generalists alike. It attracts a fair amount of attention from photographers and figures within the 20 hoverfly species most frequently recorded by photographers. It is also known to be very responsive to the effects of drought – a feature that Stuart and I drew attention to in a presentation to one of the hoverfly symposia several years ago. This relationship had previously been highlighted (but not recognised) by reports of its abundance dating back to 1947. In really hot years, the second generation is largely absent. This can be seen from past records, but sadly we generally get insufficient records to do a great deal with the data.

In 2014 R. campestris was frequently reported in April and May, and it is clear from the data that this is one of a suite of species that definitely respond to warm springs. What has happened since is more puzzling. The numbers of records have tailed off but have not dropped to a clear separation between generations. Perhaps this is the effect of northern generations emerging a little later? I must look at the data in more detail to see if this is the case, but what is clear is that the overall shape of the graph for the year to date (using a five week running mean) is somewhat different to the previous three years for which sufficient photograpic records exist.

What is also very clear is that 2013, where there was a very hard winter and later spring exhibited a clear twin-peaked phenology that is less evident in other years. The data may not be robust enough to make too much of this observation, but I do wonder if a study of bivoltine species might show how such species change their emergence patterns in response to longer breeding opportunities.

This brief observation illustrates how it may be possible to generate useful and relevant information on the effects of changing climates on our wildlife. It shows that 'common' species are highly relevant to the understanding of the natural world and should encourage more recording of such species. The big question is 'how to deal with the volume of data that could be generated by a serious initiative to record common insects?' Maybe there is scope to develop ideas by MSc students?
Yearly phenology for Rhingia campestris using photographic data using a five week running mean

Thursday, 26 June 2014

Shifts in phenology

Epistrophe eligans is one of the commoner spring hoverflies. It's larvae are often predacious upon aphids on fruit trees and therefore it can be quite common in gardens. When I first started recording hovers in the 1980s I saw it most frequently in May. Bu 2000 its earliest dates were in the third week of March This year is was 9 March! It is clearly very responsive to temperature and could be a really useful model for following climate change.

This year there have been good numbers of photographic posts of this species. The majority of records are from the midlands and southern England, with far fewer records from northern England and Scotland (the latter is at the extreme of its range). I therefore wondered if I could show differences in emergence at different latitudes? So, the following graph is split along three lines - south of a line between the Severn and the Thames set by the grid squares SSS, ST, SU and TQ; midlands between the Severn-Thames line and a line between the Dee and the Humber set below the grid squares SD, SE and TA; and a final area north of the Dee-Humber line.

Even using very limited data for one year, the differences are clear when the data are cleaned by creating a three week rolling mean (Figure 1.).

Figure 1. Phenology of Epistrophe eligans in 2014 using photographic data.
I tried the same technique with my own data over 4 epochs, 1980s, 1990s, 2000-2009 and 2010 to date. This looked as though it would work too, but I found that the change in my main recording area has affected the data - my 1980s and 90s data were assembled in southern England; now it is mainly from the midlands, so the graph does not really work.

These two  simple observations help to show how a network of recorders might generate data that could be used on a yearly basis to follow the effects of inter-yearly variation in numbers. Some similar results emerge from data for Episyrphus balteatus (Figure 2). This graph is based on the data from photographers too - and again it shows how numbers can vary hugely from year to year.

Figure 2. Inter-yearly phenology of Episyrphus balteatus based on photographic records.




Sunday, 4 May 2014

Spring data




April has passed and a considerable number of photographs have been taken, creating the biggest block of photographic data for this part of the year (Figure 1) that the HRS has assembled. Two points can be recognised. Firstly, the contribution of the UK Hoverflies Facebook group has been very significant. Secondly, that numbers of records derived from Flickr and iSpt etc have held up, despite the shift towards Facebook where several of the most active photographers now post their shots
Figure 1. Five-week running means of photographic records for the years 2011 to 2014, with the 2014 records represented as all records and a sub-set from sources other than Facebook (the situation that obtained before 2014).
The two trends offer a very positive indication of growth in hoverfly recording, which is very encouraging. It helps to show that the strategy that Stuart and I embarked upon five years ago is paying dividends. At the time, we were concerned that recruitment of new recorders was pretty low and that we were not connecting with people. That point represented a watershed between the traditional recording scheme and society-based approach to engagement, and more pro-active approach using 'virtual societies' where we set out to make contact by engaging locally. Our training programme has certainly had this effect and the internet has greatly expanded the basis for recording. Translating this into the full-spectrum recording that the old guard of the HRS have done may take a while longer, but even if only 5% of enthusiasts buy a microscope and start to look at preserved specimens we should replace the old guard with a new generation!

I think that the community of photographic recorders may also be extremely helpful in developing knowledge about yearly changes in phenology amongst more abundant and obvious species. This could be extremely interesting if developed as long time-series data. By way of example, two graphs are indicative - the first being a selection of spring species (Figure 2) and the second specifically Epistrophe eligans in 2014 (Figure 3).

Figure 2. Five week running means of combined photographic data 2011-2014 for a selection of spring hoverflies.
Figure 3. Numbers of photographic records of Epistrophe eligans in spring 2014 - data to 4 May and consequently incomplete.
The  data for spring species are combined for 2011-2014 and have been smoothed by a five week running mean. I will look in more detail at individual year data but everything is a bit skewed by larger numbers of records from 2014 and much more limited numbers in some years. I have simply posted the data for 2014 for Epistrophe eligans because in past years this species has not figured greatly in photos. The graph will be updated at some point because this species' flight season is not over. The trend is obvious though!


Sunday, 20 April 2014

What flies do photographers focus on?


As interest grows in the possible uses of web-based recording, it crossed my mind that it would be worth looking at the data for flies. Until last year I did not maintain a log of hoverflies that could not be identified, so the dataset prior to mid-summer 2013 is only partial. I therefore selected the data for 2014 only, working on the principle that these data would offer a reasonably representative example of the situation between January and April. Obviously there will be much greater scope for evaluation at the end of the year, especially as there will be lots more species recorded during the summer.

Nevertheless, the data paint an interesting picture. They comprise all photographs where reliable locality data can be attributed. The only exception is that I do not extract records of tachinidae from iSpot because I know that Matt Smith (Tachinidae Recording Scheme) does this. However, I suspect this will make relatively little difference as the numbers of tachinids posted thus far have been quite small.

The graph (Figure 1) tells a facinating story. Firstly, it is clear that hoverflies make up the vast bulk of records. This year, the numbers are perhaps skewed by the development of the Facebook page (an additional 263 records over and above data from other sources), so I have excluded these records from the graph. The other families that attract attention are the bee-flies (Bombylius major dominates but there have been a few B. discolor), the Calliphoridae (bluebottles etc), what I will call 'Muscoidea' because this tends to be a dump for bristly jobs that I cannot readily place to family, and the Scathophagidae (dominated by Scathophaga stercoraria but with several Norellia spinipes). Other families figure quite lightly at the moment but I expect the Bibionidae to surge forward over the coming weeks.

Figure 1. Numbers of photographs of Diptera families logged for 2014 (1 January to 20 April 2014).
 
Why is there such a bias in the numbers? Well, clearly my focussing on hoverflies is a factor. I do not chase up shots of other flies that lack locality data, unless they are identifiable. So, there would otherwise be a bigger dataset. But, in reality, the Flickr groups and sets are not heavily dominated by non-Syrphid flies. Hoverflies are genuinely a major component of what is noticed and photographed. There are a few surprises, however; not least the frequency with which moth flies (Psychodidae) are depicted. I guess they are obliging and don't generally fly at the first provocation?

Now, having recognised and identified possible bias in the dataset, I wonder if there are other factors at work? I wonder whether there is an element of inbuilt bias because photographers start to know the flies that they or others are most able to put names to? Clearly this is not wholly the case because lots of shots of bluebottles and muscids are posted. Perhaps, therefore, the key is the frequency with which particular subject-matter is encountered, and whether it is willing to act as a model?

Is the picture that is emerging for Diptera the same for other Orders? Whilst I browse the web I get an impession that a similar sub-sample might be seen amongst the bees and wasps, amongst beetles and spiders and true bugs. I don't know whether higher levels of identification are possible amonst these groups? Somehow I doubt it? But, can it be quantified, and are there other recording scheme organisers who keep a similar log to the one I maintain. If such logs do not exist, perhaps there is a need to develop a co-ordinated approach to data extraction from the web. I think there are definite benefits to be had from internet recording, but there are also drawbacks. The question is, are there similar drawbacks for each insect Order, or are the limitations confined to certain Orders?

These issues are worth exploring further because there appears to be a dichotomy of views amongst recorders. Some are alert to the potential of internet recording, whilst others dismiss photography as a way of getting records. Clearly, those who dismiss photography have a good reson to be sceptical but it is well worth developing a much clearer picture of what can and cannot be done. Malcolm Smart and I have looked at the Asilidae (Malcolm is currently working though last year's photos). In an analysis of the Asilidae records I developed to 2012, he managed to identify the vast majority of shots (94%), generating at least one record of 20 species (74% of the family); but just six species formed 68% of the data.

I hope to develop further analyses over time, so that we have at least a basic understanding of the flies that are recorded by photography. It might take a little while, however!

Thursday, 17 April 2014

Changes in recorder effort



The advent of digital photography and of web-based data capture techniques might suggest that biological recording is entering a new and vibrant phase. I'm sure it is, if one simply looks at the numbers of records that find their way onto databases. The big challenge is to establish whether these data actually reflect an increase in recording or a change in the ways in which plants and animals are recorded?

My interest in this stems from five years trawling through photographs on a wide variety of websites. I check between thirty and fifty sites each day, and during the summer months add between fifty and one hundred new records. I've posted some of the initial results in previous reports. My suspicion is that there will be considerable differences in the nature of recording, depending upon the ease with which animals and plants are both found and identified.

My instincts say that for those species that are easy to record there has been an increase in records. In the case of groups like macro-moths ,where moth traps have gained huge popularity, the volume of records may well have increased sharply. It would be interesting to see precisely what has happened. In the case of macro moths, perhaps there is an increasingly robust dataset? Moths are pretty docile and allow themselves to be photographed. Moreover, I suspect most can be identified from photos (I may well be wrong for a few genera).

Moving on to taxa where there is a need for specialist fieldcraft, and collecting specimens for microscopic identification, I suspect that what we are seeing is a shift in the robustness of the data. This was well demonstrated by Matt Smith and Chris Raper in their presentation of the Tachinid Recording Scheme at the last Dipterists Forum AGM. Chris and Matt showed that the data set was starting to change in composition as more photographic records were acquired. The most abundant species now include several that would be far less dominant if the data were strictly from what I would describe as taxonomist sources (i.e. people who actually study tachinids in a meaningful way).

The Hoverfly Recording Scheme has previously shown how the records of  more difficult genera are declining as a proportion of the dataset. This was apparent well before 2011 when we published the last atlas and continues to follow a similar downward trajectory (Figure 1). If anything, I suspect the rate of decline may be increasing. The big question then arises as to the cause. Is it simply that there are more records but no greater number of taxonomic recorders? Or, has there been a demographic shift?


Figure 1. Trend in the relative proportion of difficult species represented in the HRS dataset
To investigate this, I started to look at the composition of HRS recorders. We have previously noted that upwards of 50% of the data were supplied by just 21 recorders! The actual proportions bounce about quite a bit on a yearly basis so I looked at the contributions made by the top ten and top 20 recorders in any one year (Figure 2). This tells a fascinating story. All 20 recorders started to contribute between 1976 and 1992. In other words, the longest serving have been contributing for nearly 40 years and the youngest for 22 years! There were two big influxes, one in 1976/7 and a further recruitment in 1983-1985. Thus, the bulk of the major contributors to the scheme have been doing so for 30 years or more. Many on these recorders contribute between 500 and 1000 records per year, as shown in Figure 3 which depicts the contributions of the top 5 recorders. This figure shows how contributors data fluctuate quite substantially, reflecting rising and waning interest or the natural pressures of life.

Figure 2. Contributions to the HRS by the top 20 recorders over the period 1976-2013

Figure 3. Levels of data contributions by the top five contributors to the HRS. RM is 'Recorder 2'

If, however, one looks at the top 20 contributors each decade from 1976 to 2013 (i.e., 1976-1980; 1981-1990; 1991-2000; 2001 - 2010; and 2011 - 2013) the picture is very different (Figure 4). True, the same principle names are there, but the graph shows how major recorders only make a major contribution for a relatively short period of time. Critically, however, those recorders who started in the 1970's and 1980s have continued to form the nucleus of the Recording Scheme for much of the time. The data for the period 2011 to 2013 are misleading because there are several datasets that have not been updated since the call for records in 2009/10. Even so, the numbers of active new recorders are very encouraging and include at least 10 alumni of the training courses we have run in the past five years. This is most encouraging because the courses were intended to recruit replacements for the 1970s and 1980s cohorts and appears to be doing precisely that.

Figure 4. Contributions by recorders according to decade when they first became a major recorder.
The bigger question remains as to any changes in the types of data we are receiving. I think we can say with some confidence that the historic foundations of the scheme are slipping away. These were people who used Stubbs and Falk, and generally did all taxa. Many of the new recruits are equally well trained and also use Stubbs and Falk across all taxa. But, it seems likely that the absolute numbers of records entering the scheme has dropped in recent years because although new recorders have been recruited, the older ones who provided big blocks of data are less active. We can see this from Figure 5. It therefore seems highly likely that the composition of the dataset is changing, at least in the short-term. Its prognosis for the future is less clear, because we do seem to be recruiting good numbers of people who are versed in the skills needed to deal with difficult genera.

Figure 5. Numbers of records submitted to the HRS by conventional means, with the added contribution from photographs.