Showing posts with label database right. Show all posts
Showing posts with label database right. Show all posts

Thursday, 20 October 2011

Can we force facebook to give us its "like" database?

Jim Killock of the Open Rights Group pointed me at an interesting response made by facebook to an Irish student named Max's subject access request under Irish data protection legislation which forms a part of the Europe versus facebook campaign.

The particular point that interests me is that, concerns facebook's tracking of all pages visited which show a "like" button - a practice that can be really intrusive. Although Max did obtain a considerable quantity of information, facebook did not release to him their list of "like" tracked data.

In their response facebook say:

Section 4(12) of the Acts carves out an exception to subject access requests where the disclosures in response would adversely affect trade secrets or intellectual property. We have not provided any information to you which is a trade secret or intellectual property of Facebook Ireland Limited or its licensors.

Unfortunately for facebook, that isn't quite what the relevant Irish legislation appears to say (health warning: I am not an Irish lawyer). What section 4(12) of the Irish Data Protection Act 1988 says, according to a consolidated version of the statute, is:

(12) Subsection (1)(a)(iv) of this section is not to be regarded as requiring the provision of information as to the logic involved in the taking of a decision if and to the extent only that  such provision would adversely affect trade secrets or intellectual property (in particular any  copyright protecting computer software).

Note the phrase "information as to the logic involved in the taking of a decision". What this is all about is that section 4 gives data subject several different rights. One right (in section 4(1)(iii)(I)) is to be supplied with a copy of "the information constituting any personal data of which that individual is the data subject" - a simple right to information. Another, and different right, can be found in section 4(1)(iv) which applies to automatic decision making by the data controller. Here the data subject has a right to be informed of the "logic involve in the processing". Obviously that's quite a different right since it is essentially a right to know about algorithms rather than data.

Quite clearly section 4(12) is a restriction on the right under 4(1)(iv) to know about the logic of automatic decision making and not a restriction on the right of information simplicter. Nice try facebook, but I can't see that working.

Our own legislation is very slightly different. We also have a right (in section 7(1)(d) of the Data Protection Act 1998 to be informed about the logic involved in automatic decision making, but the restriction on that right is limited to trade secrets. Section 8(5) says:

Section 7(1)(d) is not to be regarded as requiring the provision of information as to the logic involved in any decision-taking if, and to the extent that, the information constitutes a trade secret.
So that any UK national involved in the Europe v facebook campaign has a much stronger argument.

In any case, at best facebook can claim a database right over the contents of the list of pages visited by Max that they have collected using the "like" button. The database right is a creature of European law (directive 96/9/EC). Recital 48 of the directive states that "the provisions of this Directive are without prejudice to data protection legislation", which seems to me to argue that data protection law ought to trump database right. If you think about it, the contrary would be an impossible situation. Personal data will often be protected by database rights. If you could use database rights to avoid subject access requests they would be of far less use.

Thursday, 28 October 2010

Database right: proving infringement

A recent decision in the Patents County CourtBeechwood House Publishing v Guardian Products [2010] EWPCC 12 concerns the database right. We haven't seen very many database right cases, so I thought it was worth a short comment.

The claimant, who I shall call by their trading name "Binley's" maintain and sell a database of the names and addresses of people associated with GP practices (such as doctors and nurses) which is then sold to companies that wish to use the addresses for direct marketing. It costs, according to Binley's, roughly £110,000 a year to keep the database up to date.

In order to detect any infringement of their database right, Binley's include in their database a number of what they call "seeds". These are bogus entries, giving the address of Binley's staff. When any post is received addressed to a seed address, Binley's can then presumably check the source against their list of clients to check that the marketing comes from someone authorised to use their database.

In about August 2007 Binley's received a letter addressed to a seed. The letter was from Guardian Products (the first defendant) who had obtained their mailing database from the second defendant (Precision Direct Marketing Ltd) who had in turn obtain it from an organisation called Bespoke Database Organisation Ltd (or BDOL for short). Quite why BDOL were not also defendants is unclear. Perhaps they had already reached a settlement with Binley's but we do not know. But it was accepted that BDOL's database contained the offending seed.

Binley's sued the defendants for infringement of their database right. Presumably on the ground that the defendants must have "extracted" — which means (in English law at least) "the permanent or temporary transfer of [the contents of the database] ... to another medium by any means or in any form" (see regulation 12 of the Copyright and Rights in Databases Regulations 1997).

One assumes that Binley's felt their case was pretty strong, so they made an application for summary judgment. A summary judgment application is not a trial, to succeed Binley's needed to persuade the judge that the defendants had "no real prospect" of defending the claim.

On this point Binley's failed. The problem was, from the judge's point of view, that all the evidence he had was:

  • Binley's database contained one or more seeds
  • One of those seeds had turned up in BDOL's database
Clearly this was evidence of some copying, but extraction from a database is only unlawful if the extraction is all or a substantial part of the contents of the database (see regulation 16). It might seem highly improbable that the offending seed was the only item copied. Indeed the judge thought that it was "highly probable" the there had been the extraction of a substantial part, that that was not enough to grant summary judgment.

Binley's had not given evidence of the proportion of seeds in their database. Its evidence was there were "a few" but Binley's had refused to give a more precise figure, possibly for commercial reasons. If they had, that might have allowed some assessment of the degree of extraction and therefore whether it was not substantial.

Summary judgment was refused.

What is interesting about this case is that it may be very difficult in practice to prove merely by examining a database's contents that it has been copied from another, especially where the data is relatively regular and commonplace such as names and addresses, without recourse to seeds or some other form of watermarking. The more seeds or watermarking, the easier the task, but at the cost of poisoning the database owner's product with irrelevant or false information.

In some cases other evidence will be available that demonstrates extraction, but here the first defendant appears to have had no direct knowledge of how its database had been created (since it originated in BDOL). In such a situation a database owner may have difficulty proving that the extraction of a substantial part has taken place.

Note that a decision of the Patents County Court sets no precedent. Its value is merely illustrative, but I feel it is interesting nonetheless.

Tuesday, 4 August 2009

National Portrait Gallery: is there a database right?

I have already written about the National Portrait Gallery's legal threat against Mr Coetzee, an editor of wikipedia. I considered only the validity of the gallery's copyright claim. What about its claim that Mr Coetzee infringed the gallery's database right?

The database right is a based on the EC Directive 96/9/EC, transposed into English law by the Copyright and Rights in Databases Regulations 1997.

The directive defines a "database" as:
a collection of independent works, data or other materials arranged in a systematic or methodical way and individually accessible by electronic or other means
That's a pretty broad definition and covers anything you might think of as a "database". The photographs of paintings are independent works (even if they are not subject to copyright protection) and the gallery seem to have arranged them in a systematic way so that they are accessible by electronic means. That just tells us what a "database" is. In order to obtain the protection of the database right, the maker of the database must show:
that there has been qualitatively and/or quantitatively a substantial investment in either the obtaining, verification or presentation of the contents to prevent extraction and/or re-utilization of the whole or of a substantial part, evaluated qualitatively and/or quantitatively, of the contents of that database
That is actually 6 different conditions, all neatly packed up in the compression system that is legal drafting. First the maker of the database must show that there has been some substantial investment, which can be of two kinds, either quantitative or qualitative; and second that investment can be in one of three things: obtaining, verification or presentation.

What sort of investment counts? Here the decision of the European Court of Justice in Fixtures Marketing v OPAP C-444/02 comes into play. The top divisions in English and Scottish football drew up fixture lists for the matches to be played in the various divisions during each season. Fixtures Marketing Limited had been assigned the rights to manage this information outside the United Kingdom. OPAP repeatedly extracted the names of pairs of football teams playing against each other and displayed them on its website. Fixture Marketing were obviously unhappy about this and sued. The question of what was a database and how was it infringed was referred to the European Court of Justice.

The court held that "investment in ... the obtaining ... of the contents" referred to resources used to seek out independent materials that already existed and to collect them into the database and not the resources used to create the independent materials in the first place. The purpose of the database right was (thought the court) not to protect the creation of materials that went into the database but to protect the creation of the storage and processing systems for the database.

As support for this view the court pointed out that recital 19 of the directive states that the compilation of several recordings of musical performances on a CD does not represent a substantial enough investment to be eligible. The investment in the music is not enough.

The relevance of this case to the NPG is obvious. One suspects a great deal of the investment in creating the photograph database was in taking and making the photographs. That, on the authority of Fixtures Marketing is irrelevant.

As for verification and presentation, the court seems to have required that any investment that forms a part of the creation of the independent materials would also have to be disregarded. For example, on verification the court said:

The professional football leagues do not need to put any particular effort into monitoring the accuracy of the data on league matches when the list is made up because those leagues are directly involved in the creation of those data. The verification of the accuracy of the contents of fixture lists during the season simply involves, according to the observations made by Fixtures, adapting certain data in those lists to take account of any postponement of a match or fixture date decided on by or in collaboration with the leagues. Such verification cannot be regarded as requiring substantial investment.

Of course I have no idea exactly what work went in to the NPG's database, but I think it is here that they come unstuck. In their letter to Mr Coetze the gallery's solicitors use the time honoured tactic of proof by assertion:

Our client’s website includes a searchable database of over 60,000 carefully chosen, curated and watermarked images. There can therefore be no doubt that our client’s database of images is a “database” for the purposes of s.3(A)(1) of the CDPA.

If, as seems to be the case, the gallery had set aside funds to digitise its collection, those funds and the investment they represent would have to be ignored. The "carefully chosen" images would be just those images that the gallery had chosen to digitise, creating a database from that collection would not necessarily represent the substantial investment the directive envisages. I have no idea whether the "curated" refers to the paintings (in which case it is irrelevant) or the images (in which case I don't know what curating a digital image might mean). The watermarking would surely form a part of the gallery's effort in creating the digital images in the first place and also have to be disregarded.

There's too little information to be able to assess how good the gallery's claim might be, but whether they have a database right at all is open to some considerable doubt.



Wednesday, 8 July 2009

Post codes and the database right

At opentech 2009 Harry Metcalfe presented the idea for a site (which I will call Ernest Marples) to convert postcodes into latitude and longitude pairs. Ingeniously, the site does not store any data itself, instead it scrapes a number of other sites for the information and returns the result. Does this get around the Royal Mail's database right in the postcode database? Sadly, I do not think it does.

Let us ignore for the moment whether the site breaches any terms and conditions of the sites that it is scraping for its data. The identity of those sites is a secret so although there amy be a question mark over the legality of the scraping, to say any more would be to theorise without data. Instead I want to focus on the database right.

In the 1990's it became clear that, in a number of EC/EU member states, that collections of information could not necessarily be protected by copyright. For example in the Dutch case of Van Daele v Romme a publisher was unable to prevent the copying of all the words in its dictionary. A similar position was reached in the United States where the Supreme Court decided in Feist that a telephone directory could not be subject to copyright.

The eventual result was Directive 96/9/EC of the European Parliament and of the Council on the legal protection of databases. Chapter III of the directive creates a thing called the "sui generis right". The core of that right can be found in article 7(1):

Member States shall provide for a right for the maker of a database which shows that there has been qualitatively and/or quantitatively a substantial investment in either the obtaining, verification or presentation of the contents to prevent extraction and/or re-utilization of the whole or of a substantial part, evaluated qualitatively and/or quantitatively, of the contents of that database.
As you can see a lot of alternatives are being packed in. If all that you remember is that the focus is substantial investment you will not go far wrong. Roughly speaking, lots of investment implies protection. To unpack a little: the substantiality of the investment can be qualitative (it took real skill to select just these poems) or quantitative (we spent many person years walking to every grid point and photographing it). That investment can be in verification and presentation as well as collection.

The right allows a rights holder to prevent either:

  • extraction; or
  • reutilisation
Of a substantial part of the database.

Ernest Marples is not extracting a substantial part of the database, but I'm less sure about re-utilisation. The term "re-utilization" is defined in article 7(2)(b) to mean:

any form of making available to the public all or a substantial part of the contents of a database by the distribution of copies, by renting, by on-line or other forms of transmission.
Is that what Ernest Marples site is doing? On the one hand Ernest Marples only hands out single (postcode, co-ordinate) pairs and so it could be argued that it is not making the whole of the database available at any one time. On the other hand a member of the public can query any postcode and Ernest Marples is almost certain to be able to return a result for it.

It may be that the drafters of the directive saw Ernest Marples coming because they added an additional form of infringement in article 7(5):

The repeated and systematic extraction and/or re-utilization of insubstantial parts of the contents of the database implying acts which conflict with a normal exploitation of that database or which unreasonably prejudice the legitimate interests of the maker of the database shall not be permitted.
Roughly speaking: lots of insubstantial extractions etc may add up to a substantial one. When exactly? That was just the question that the Swedish Supreme Court in Fixures Marketing v AB Svenska and the Court of Appeal of England and Wales in British Horseracing Board v William Hill wanted to know. The European Court of Justice explained to them that the purpose of article 7.5 is exactly to prevent someone getting around 7.1. If the effect of the repeated extractions or re-utilzations would have the same negative effect on the maker of the database as a breach of 7.1 (if all the extractions etc had been done all at once), then that is a breach of 7.5.

I think Ernest Marples is probably caught by 7.5 even if he gets away with avoiding 7.1. Article 8 does provide a defence for lawful users of the database but that is even more fraught a line of argument (if you thought 7 was badly drafted, have a read of 8 and try and work out what its meant to do). There is some doubt, but not enough to make Ernest Marples's method your business plan.

As Harry explained at opentech, part of the purpose of the site is political. There is a strong body of opinion that the Royal Mail should not have a monoploy on this extremely important database. If Ernest Marples is sued that will generate terrible publicity for the Royal Mail (as it should).