All posts by David Sewell

Going on a date with Google – Dublin Core and More

It’s all about the first date… and this post covers some of the devious lengths Google will go to to get it.

Why does Google want to get a date out of you?

Time is generally accepted as an essential ingredient in the calculations of relevancy in SERPs. Some events will be promoted in SERPs because they are news, others, such as recurring events, may well appear in results at appropriate times. Updates and edits to site content can help revive an older story. All these examples are dependent upon the date attributed to the article, the origin and the changes made. The minty freshness update back in November 2011 caused noticeable impact to around 10% searches.

To try to benefit from these changes, many site owners looked for ways to touch up older pages and content to make it at least look more current. It is the equivalent to software folks in Unix world reaching for the ‘touch’ command to change the ‘last modified dates’ on all their files to ‘get things noticed’.

 

But what is Google looking for?

With the drive to discover fresh content and attempts to improve the understanding of timelines, Google will try to ‘datestamp’ everything it finds… and it can make mistakes when generating a datestamp for your content that impacts your relevancy to search, especially if ‘date order’ is selected as an advanced search option.

 

Unique IDs as Dates

Many e-Com sites use unique identifiers in page titles in order to reduce duplications and create unique, informative pages. Readers are likely to be familiar with GWMT diagnostics section on HTML suggestions:

Non Informative Title Tags and duplications

 

In order to reduce the number of duplicated page titles, a typical workaround would be to use the following format for product details pages:

{brand} {product} {key feature} – {unique identifier}

The purpose of the unique ID is to ensure that varieties of the same product (e.g. sizes, colours) each have a unique page title. [You may have noticed GWMT also reports ‘noninformative’ title tags… strongly suggesting that the title remains a key source of information for Google’s indexing algorithms]

For an average SEO making page title unique is just the start. Introducing an ID can guarantee uniqueness and introduce useful information (win-win!), but it can have unintended consequences (oops!).

The purpose of this post is to cover just one of these unintended consequences – misinterpreting a product ID for a datestamp.

The search result below shows an example of Google misinterpreting a product ID and using it as a datestamp in results:

Product ID used as date stampThe unique identifier 11171977 has been interpreted as a date. Now this listing carries the datestamp: 17 Nov 1977 which has been taken from the product ID (11-17-1977)

That’s inconvenient if you’re trying to sell this house, it looks like its been on the market since 1977 !

Here’s another example (pink camera)- this time it’s the model number that is being treated as a date:

Pink Camera product ID as date stampThe DSC-T70 8.1 in the title is being treated as if it means 1st August 1970 !

Dates in Content

When an article is posted on a blog, there is often a time-stamp associated with the post and prominently displayed at the top of the page. Google finds this on the page and the association is made, in this case, quite sensibly assuming your server time is correct and you aren’t manually fiddling the dates.

Less obviously, if you have a table or other date source within your content, there is a chance Google will take that date to be an appropriate datestamp. For instance, if you reference the last time a certain event happened, but fail to state the current date, then the entire article could be catalogued as the time of the referenced event rather than the current one.

Be careful to check your content for date references and help Google find an appropriate date for your content.

If you see a strange datestamp in SERPs, then look at these on-page elements for clues as to the cause of the mix-up:

a) Footer for software version numbers

b) Tables of reference data

c) References to dates within text (ensure the first date referenced is the date you’d like to have associated with the post)

d) Telephone numbers

 

Avoiding a bad first date:

Here are a few ways you can help make your first date with search engines more memorable:

a) Place the preferred date between the article’s title and the article’s text in a separate line of HTML, and remove all other date references from the body of the article to avoid confusions.

b) Clearly mark up the datestamp you want associated with the article using the Dublin Core meta tag:

<code><meta name="DC.date.issued" content="YYYY-MM-DD"></code>

or more completely:

<code><link rel="schema.DCTERMS" href="http://purl.org/dc/terms/" />
<meta name="DCTERMS.created" content="YYYY-MM-DD" />
<meta name="DCTERMS.modified" content="YYYY-MM-DD" />
</code>

c) Use the <publication_date> tag inside news sitemap xml files

d) If using OpenGraph, then add the article:published_time, article:modified_time tags

e) Using microformats and schema mark-up to ensure telephone numbers are parsed correctly (see Schema.org for these definitions), ie. not treated as dates.

f) Modify the way product IDs are displayed. Introduce characters or spaces to break up the ID so it becomes less ‘date-like’, yet retains uniqueness.

If erroneous Google datestamps have caused you problems with search relevancy, please let me know !

 

Enhanced by Zemanta

301 redirect URLs reappearing in Google – Phantom menace

Over the last few months, phantom 301 URLs have been showing up in Google search results.

What is a phantom 301 URL?

These are old URLs that have been previously 301d and dropped from the Google index, but which are now showing up again. (Read: 301 redirected showing up again after 1 year)

Example 1:

PropertyFinder was bought by Zoopla and 301 redirected correctly with a link map that matched appropriate pages to each other across domains.

Until recently, the PropertyFinder website would not show up for the search, as the site had been ‘de-indexed’ in Google. However, looking at the screenshot below (Jan 2012), it is clear that PropertyFinder now has a Phantom URL in first place in Google search results:

Phantom 301 URL

 

Example 2:

Dothomes.co.uk was bought by the DPG Group and redirected to FindaProperty.com

Again, as with PropertyFinder, the DotHomes website would not show up in results until recently. Performing a site: search for this old domain suggests Google has acknowledged the permanent redirect:

Google acknowledges the 301 redirect

However performing a direct search for the old domain (dothomes.co.uk) shows a mixed up result that:

  • displays the old 301d URL (dothomes.co.uk/)
  • uses the old 301d URL as destination for the current FindaProperty.com title
  • the current FindaProperty.com description and sitelinks

DotHomes shows in search results

 

Phantom 301s have been appearing more frequently recently (since Dec. 2011), and understanding the full implications of this change in Google search results is important if:

1. You have a website founded on a hotchpotch of old websites that have been 301’d

2. You are considering a redirection project, such as conversion to friendly URLs from unfriendly, parameter based URLs

 

What is the impact of a phantom 301?

Websites that have been bought out, or taken over and had their pages correctly 301 redirected to new sites several years ago are most affected by this change.

The historic 301d URLs from these old sites are re-appearing in search results, taking up positions they used to occupy before the 301 was put in place. The impact is at least threefold:

  • There is a new duplicate content risk for the page redirected to (as the redirection has ‘failed’)
  • The search results show an old entry, with a less optimised call to action (historic titles, URL & snippet)
  • Equity may not be passed to the destination URL (as the redirection has ‘failed’)

Before you say, ‘I know how 301s work, and what you are saying simply isn’t true’, read on!

How 301s used to work in the ideal world

What used to happen, is that Google would crawl the old URL, receive the 301 header response and new location, then pass ‘some’ link equity to the new URL, index it and de-index the old URL over a period of a few days.
The upshot is that your new URL would rank immediately, and the old URL would be gone forever…

What’s changed ? Is this Panda related ?

What has started to happen, is these old URLs have popped back into SERPs.
There have been no technical changes, these old URLs are still 301’d (no changes made for many months / years in some cases), but the they are being presented in SERPs carrying the old site name and titles, showing the old URLs and often having much lower quality CTAs.

So the smoking gun points to Google tweaking the algorithm…

It is as if 301 redirects are no longer guaranteed to remove a URL from the index.

If Google believes the old URL is the best result for the search, then it will show it.

Digression follows:

My personal hunch, is that Panda is strongly related to CTR testing and bounce rates. Sites that have suffered under Panda have tended to have the sort of content you’d ignore (packed with adverts, poor layout and spelling etc), or the URL in the search results that you’d usually skip over, because the last time you clicked that site, it was full of articles based entirely on spun content.

Panda nailed those sites, they were the sites people didn’t like to click on, or stay on.

The phantom 301 URLs could be the latest extension of presenting and testing the URLs that people prefer to click. They are back in the popularity contest, despite the permanent redirect. An old URL may have had the best CTR. The only way to find out, is to re-present the old URL in the search results and see how it fares. If it receives a lot of clicks, then the phantom menace may be there to stay, whether it is 301d or not.
This change may leave Google open to abuse and may change the way we work with old domains

Here are some ideas:

  • Create a lot of duplicate pages with poor titles and CTAs, then 301 them to competitor URLs in the hope that some of the ‘fresh’ pages create phantoms and devalue the competitors original URL.
  • Occupy more search positions by creating multiple URLs within a site (such as changing to friendly URLs), to allow ranking of the old unfriendly URL that has been 301d alongside the new friendly URL. Continue to add new URLs and 301 to old URLs to create a 301 farm. Google may show some of the new URLs alongside their duplicates for CTR testing.
  • Link build to old domains (if they were popular) that currently 301 to your site. These old domains are often recalled strongly and continue to be entered in search results, even when a company’s name has changed. For example, even though Santander bought Abbey, searching for ‘Abbey National’ still has a lot of suggestions, and shows phantom 301 results in the SERPs:
Abbey National Phantom search result

If you have any experience of this phantom 301 phenomenon and how it could be used to advantage, please get in touch.

Enhanced by Zemanta

fb_xd_fragment added to URLs with IE 7.0

Seeing URLs with fb_xd_fragment= reported in Google Analytics may not mean that your SEO is suffering, so don’t panic!

Your alarm might rise further when you see that the presence of the fb_xd_fragment= parameter causes the entire page content to become hidden whilst the page is loading. Still, don’t panic – investigate!

There are three main situations when panic is justifiable:

1. If you see that Google has indexed pages from your site with the additional &fb_xd_fragment= parameter, (Hint: use the site: and inurl: operators to check)

2. Users are reporting that they are seeing ‘blank’ pages on your site e.g. http://www.letsrun.com/forum/flat_read.php?thread=4349142 If you have come to this page looking for a solution as a user of a site, try to delete the ‘&fb_xd_fragment=‘ from the end of the URL and reload the page. if that does let you see the page then contact the webmaster and direct them to this post so they can implement a fix.

3. Since the confirmation that Google is indexing facebook comments (cf. Matt Cutts http://www.searchenginejournal.com/google-indexing-facebook-comments/35594/), the indexation of bad URLs from your site may grow as a consequence. Google may read the javascript on your site and generating these URLs to find additional pages to crawl.

Why are URL including &fb_xd_fragment= being visited in the first place?

The standard implementation of Facebook Javascript SDK for ‘Like’ buttons has been cited as the cause of fb_xd_fragment being appended to URLs for users of “certain browsers”.

It appears that Facebook Javascript SDK causes these ‘phantom’ visits when real visitor clicks on a ‘Like’ button (and overwhelmingly, the visits come from visitors using IE 7.0). It is a combination of the XFBML version of facebook plugins and IE 7 that gives rise to these URLs being generated by javascript and then reported in GA.

Now for a run through of the solutions to the fb_xd_fragment bug…

Here’s a recipe for various solutions depending upon the scale of your problem:

fb_xd_fragment solution A:

Use a custom channelURL file (channel.html) stored at root level. This is used to overload the FB.init() constructor by setting the channelUrl parameter value when FB initialises.The contents of channel.html are simply:

<script src="//connect.facebook.net/en_GB/all.js"></script>

The file above must be cached for as long as possible to provide a smooth user experience and avoid time delays.

This file can fix cross-site scripting issues (see http://developers.facebook.com/docs/reference/javascript/) – Facebook dev resources refer to this issue with tongue in cheek: “The channel file addresses some issues with cross domain communication in certain browsers.”

Finally, add the following code to your pages (on detection of IE 7.0):

<div id="fb-root"></div>
<script>
window.fbAsyncInit = function() {
   FB.init({appId  : 'MY APP ID',  status : true,  cookie : true, xfbml  : true, 
            channelUrl  : 'http://www.yourdomain.com/channel.html'});
};

(function() {
var e = document.createElement('script');
e.src = document.location.protocol + '//connect.facebook.net/en_GB/all.js';
e.async = true;
document.getElementById('fb-root').appendChild(e);
}());
</script>

 

fb_xd_fragment solution B:

Use javascript to undo the nasty side effects of the fb_xd_fragment= in order to make the hidden html visible once again. Use onLoad() to run the following script. NB. As with all scripts, this solution may not work every time, as is essentially a sticking plaster to undo what has already been done.

<!-- Correct fb_xd_fragment Bug Start -->
<script>
document.getElementsByTagName('html')[0].style.display='block';
</script>
<!-- Correct fb_xd_fragment Bug End -->

 

fb_xd_fragment solution C:

For those of you that would rather not use a custom FB.init and channelUrl, or add script to undo this ‘blanking’ effect, then the cheaper option is to use a redirect. Use 301 redirection to ensure that any URL requested with the fb_xd_fragment is redirected to the same URL only without this fragment.

 

fb_xd_fragment solution D:

Taking another step back from addressing the cause, you could try to prevent indexing of these pages by using the rel canonical link. Ensure that rel canonical is defined without the fb_xd_fragment on all pages that have Like buttons (use Browser/User Agent detection to set the canonical to keep page weight down, if preferred).

fb_xd_fragment solution E:

Use WebmasterTools Site Configuration / URL Parameter settings to set to ‘ignore’ the fb_xd_fragment parameter. This informs Google that URLs with this parameter are not important, and should not be indexed.

fb_xd_fragment solution F:

Fix (or more accurately, ‘hide’) the reporting. To keep these ‘phantom’ additional page views out of GA, add a filter in Google Analytics to strip all pages that include the fb_xd_fragment. Ensure that you have identified and fixed (if necessary) the cause before filtering these pages from GA! If you have implemented Solution A, then you don’t need to use a filter. The problem of reporting the ‘phantom’ pages should go away with Solution A!

 

As with most SEO work, pick the solution appropriate to your needs.

Footnote:

If you see visitors appearing to come from a Google cache with this parameter, it is likely that these visitors use iGoogle and they are following a cached link from their dashboard. Look at the rlz parameter in the cache string for G1 to identify the source as iGoogle:

iGoogle cache with Internet Explorer 7 rlz parameter

 

Enhanced by Zemanta