Future or No-Future?
Now that's a headline I had no clue that I'd ever use! Too dramatic, maybe? I'll let you be the judge. Apparently, the future is “A.I.” and there is nothing to be fearful of…
Even the US President is boasting openly about it, and no figure he ever mentions seems to have fewer than ten trailing zeros. No bad thing when it comes to medical research, I'd adjoin with. However, the real power behind these megalithic goals is simply information, aka data, more often than not, everyone else's…
The are at least two weaknesses. One is in Webmasters like myself not being able to protect their information from such predatory theft. Another is that no real check is made on the veracity of any information that is stolen. So the old “GIGO” adage will become more true than ever!
Up to now, I've been confident that Copyright laws are strong enough to withstand most thefts of information. However, what if that was no longer true? From what I've experienced recently, I believe that the days of this being a meaningful protection may have passed.
Open Season
What's happening is that harvesting robots are scraping their “data” from everywhere they can. Read scraping as stealing everything in sight, totally flouting any basic rules that Webmasters set for well behaved robots. We mostly want to help Web robots because they index our content and usually bring new visitors in their wake. These parasites do nothing but grab, without so much as an acknowledgement!
Because it's a future that, mostly we don't understand, these thieves, for that is what they are, feel entitled to harvest any and all information they can get their grubby mits onto. Oh, they're totally brazen about it! So, it's open season for attacks on any and every Web site out there.
You might think, so what? Not content with flouting all accepted standards, the pesky blighters have issued millions of prying spies. The Timeline has been adversely hit every four seconds by them. What's even more worrying is that many of these pests employ “advanced technology” so they look for and take everything!
From China came a robot that took 58 pages of information and used over 12Mb of bandwidth in just seconds. As a human visitor, you would really struggle to download anywhere near that amount of bandwidth from such a small page count.
From what I can glean, the stolen information is stored, then processed in advanced databases. Eventually it will be presented as if the thieves were the originators. I seriously doubt that it will be free to obtain or view at that point.
Of course, all the work to display the information in a readable fashion will have already been done, by me, before their raid. So they're also making use of the fact that all the Timeline pages validate to W3C standards. This just adds insult to injury.
Powerless
That's legalised plagiarism, I hear you say. Well, yes it is, however it's not at all legal! The problem being, where am I going to find sufficient funds to fend off these mega-buck pirates? At one time Google was the best friend of the small guy like myself.
Sadly, it seems that they might now be leading this tussle. Their robots have not always been well behaved and certainly not obeying limits. The attitude now taken elsewhere being that if Google can get away with it, so can they.
So, what's the impact of all this? Well, the feeling of being utterly powerless is extended to cover what is likely to happen to over 25 years of painstaking research. I won't dwell on costs, because interactions with visitors more than compensate for any financial outlay I have made.
As you probably know, my approach to recording information has been as meticulous as I can manage. There are bound to be errors, that's a given, but once the information leaves these pages those errors are baked in. It just adds to the undermining aspects.
The biggest question is: Where do I go from here? The odds are completely against me. For example, I'm at a point of probably purchasing the 1958 crew records from the National Archive at Kew. No small outlay, with not much change out of £800. This gives me the opportunity to take the crew lists forward by 2 years from where I once thought they had ended.
Having got quite good at this transcription thing, I was looking forward to the challenges they present. No matter that it's many hours of transcription work, there's a satisfying achievement at the end. However, now I need to consider whether it's worth doing, only for some pesky blighters to not only steal the work, but also go on to claim it as their own.
A key aspect of this is, that when it comes to the archive content, no-one in “officialdom” realises that the surviving official records do not record the full names of the crew, only their family name and initial(s). When the records come to me, having built so much of a back-catalogue, I'm able to match a family name and occupation, and where they have previously travelled, insert absent full names. This instead of perpetuating just family name and initials. Most “people” searches on the Web involve full names.
The Impacts
Let's start with devastating! I cannot bear liars and thieves. Their actions always hit you so personally, and of course, they are oblivious to any impact. In my study there are over 35 foolscap file boxes, mostly filled to capacity. The Timeline, not only being the collection catalogue but a source of added information that has grown and grown and, I believed, added value to the ephemera.
I have laboured under the belief that one would always accompany the other as a valuable social history archive, especially as the Timeline records so many first-hand experiences from so many sources. Now I'm wondering whether any of it has any value. I'm even wondering whether sharing this information was a mistake, though of course, I'd no inkling that it might lead to this outcome.


