Thursday, November 20, 2014

Building a complete Tweet index by Yi Zhuang

Tuesday, November 18, 2014

Today, we are pleased to announce that Twitter now indexes every public Tweet since 2006.
Since that first simple Tweet over eight years ago, hundreds of billions of Tweets have captured everyday human experiences and major historical events. Our search engine excelled at surfacing breaking news and events in real time, and our search index infrastructure reflected this strong emphasis on recency. But our long-standing goal has been to let people search through every Tweet ever published.
This new infrastructure enables many use cases, providing comprehensive results for entire TV and sports seasons, conferences (#TEDGlobal), industry discussions (#MobilePayments), places, businesses and long-lived hashtag conversations across topics, such as #JapanEarthquake, #Election2012, #ScotlandDecides, #HongKong,#Ferguson and many more. This change will be rolling out to users over the next few days.
In this post, we describe how we built a search service that efficiently indexes roughly half a trillion documents and serves queries with an average latency of under 100ms.
The most important factors in our design were:
  • Modularity: Twitter already had a real-time index (an inverted index containing about a week’s worth of recent Tweets). We shared source code and tests between the two indices where possible, which created a cleaner system in less time.
  • Scalability: The full index is more than 100 times larger than our real-time index and grows by several billion Tweets a week. Our fixed-size real-time index clusters are non-trivial to expand; adding capacity requires re-partitioning and significant operational overhead. We needed a system that expands in place gracefully.
  • Cost effectiveness: Our real-time index is fully stored in RAM for low latency and fast updates. However, using the same RAM technology for the full index would have been prohibitively expensive.
  • Simple interface: Partitioning is unavoidable at this scale. But we wanted a simple interface that hides the underlying partitions so that internal clients can treat the cluster as a single endpoint.
  • Incremental development: The goal of “indexing every Tweet” was not achieved in one quarter. The full index builds on previous foundational projects. In 2012, we built a small historical index of approximately 2 billion top Tweets, developing an offline data aggregation and preprocessing pipeline. In 2013, we expanded that index by an order of magnitude, evaluating and tuning SSD performance. In 2014, we built the full index with a multi-tier architecture, focusing on scalability and operability.
Overview
The system consists 4 main parts: a batched data aggregation and preprocess pipeline; an inverted index builder; Earlybird shards; and Earlybird roots. Read on for a high-level overview of each component.
Batched data aggregation and preprocessing
The ingestion pipeline for our real-time index processes individual Tweets one at a time. In contrast, the full index uses a batch processing pipeline, where each batch is a day of Tweets. We wanted our offline batch processing jobs to share as much code as possible with our real-time ingestion pipeline, while still remaining efficient.
To do this, we packaged the relevant real-time ingestion code into Pig User-Defined Functions so that we could reuse it in Pig jobs (soon, moving to Scalding), and created a pipeline of Hadoop jobs to aggregate data and preprocess Tweets on Hadoop. The pipeline is shown in this diagram:

The daily data aggregation and preprocess pipeline consists of these components:
  • Engagement aggregator: Counts the number of engagements for each Tweet in a given day. These engagement counts are used later as an input in scoring each Tweet.
  • Aggregation: Joins multiple data sources together based on Tweet ID.
  • Ingestion: Performs different types of preprocessing – language identification, tokenization, text feature extraction, URL resolution and more.
  • Scorer: Computes a score based on features extracted during Ingestion. For the smaller historical indices, this score determined which Tweets were selected into the index.
  • Partitioner: Divides the data into smaller chunks through our hashing algorithm. The final output is stored into HDFS.
This pipeline was designed to run against a single day of Tweets. We set up the pipeline to run every day to process data incrementally. This setup had two main benefits. It allowed us to incrementally update the index with new data without having to fully rebuild too frequently. And because processing for each day is set up to be fully independent, the pipeline could be massively parallelizable on Hadoop. This allowed us to efficiently rebuild the full index periodically (e.g. to add new indexed fields or change tokenization)
Inverted index building
The daily data aggregation and preprocess job outputs one record per Tweet. That output is already tokenized, but not yet inverted. So our next step was to set up single-threaded, stateless inverted index builders that run on Mesos.
The inverted index builder consists of the following components:
  • Segment partitioner: Groups multiple batches of preprocessed daily Tweet data from the same partition into bundles. We call these bundles “segments.”
  • Segment indexer: Inverts each Tweet in a segment, builds an inverted index and stores the inverted index into HDFS.
The beauty of these inverted index builders is that they are very simple. They are single-threaded and stateless, and these small builders can be massively parallelized on Mesos (we have launched well over a thousand parallel builders in some cases). These inverted index builders can coordinate with each other by placing locks on ZooKeeper, which ensures that two builders don’t build the same segment. Using this approach, we rebuilt inverted indices for nearly half a trillion Tweets in only about two days (fun fact: our bottleneck is actually the Hadoop namenode).
Earlybirds shards
The inverted index builders produced hundreds of inverted index segments. These segments were then distributed to machines called Earlybirds. Since each Earlybird machine could only serve a small portion of the full Tweet corpus, we had to introduce sharding.
In the past, we distributed segments into different hosts using a hash function. This works well with our real-time index, which remains a constant size over time. However, our full index clusters needed to grow continuously.
With simple hash partitioning, expanding clusters in place involves a non-trivial amount of operational work – data needs to be shuffled around as the number of hash partitions increases. Instead, we created a two-dimensional sharding scheme to distribute index segments onto serving Earlybirds. With this two-dimensional sharding, we can expand our cluster without modifying existing hosts in the cluster:
  • Temporal sharding: The Tweet corpus was first divided into multiple time tiers.
  • Hash partitioning: Within each time tier, data was divided into partitions based on a hash function.
  • Earlybird: Within each hash partition, data was further divided into chunks called Segments. Segments were grouped together based on how many could fit on each Earlybird machine.
  • Replicas: Each Earlybird machine is replicated to increase serving capacity and resilience.
The sharding is shown in this diagram:
This setup makes cluster expansion simple:
  • To grow data capacity over time, we will add time tiers. Existing time tiers will remain unchanged. This allows us to expand the cluster in place.
  • To grow serving capacity (QPS) over time, we can add more replicas.
This setup allowed us to avoid adding hash partitions, which is non-trivial if we want to perform data shuffling without taking the cluster offline.
A larger number of Earlybird machines per cluster translates to more operational overhead. We reduced cluster size by:
  • Packing more segments onto each Earlybird (reducing hash partition count).
  • Increasing the amount of QPS each Earlybird could serve (reducing replicas).
In order to pack more segments onto each Earlybird, we needed to find a different storage medium. RAM was too expensive. Even worse, our ability to plug large amounts of RAM into each machine would have been physically limited by the number of DIMM slots per machine. SSDs were significantly less expensive ($/terabyte) than RAM. SSDs also provided much higher read/write performance compared to regular spindle disks.
However, SSDs were still orders of magnitude slower than RAM. Switching from RAM to SSD, our Earlybird QPS capacity took a major hit. To increase serving capacity, we made multiple optimizations such as tuning kernel parameters to optimize SSD performance, packing multiple DocValues fields together to reduce SSD random access, loading frequently accessed fields directly in-process and more. These optimizations are not covered in detail in this blog post.
Earlybird roots
This two-dimensional sharding addressed cluster scaling and expansion. However, we did not want API clients to have to scatter gather from the hash partitions and time tiers in order to serve a single query. To keep the client API simple, we introduced roots to abstract away the internal details of tiering and partitioning in the full index.
The roots perform a two level scatter-gather as shown in the below diagram, merging search results and term statistics histograms. This results in a simple API, and it appears to our clients that they are hitting a single endpoint. In addition, this two level merging setup allows us to perform additional optimizations, such as avoiding forwarding requests to time tiers not relevant to the search query.
Looking ahead
For now, complete results from the full index will appear in the “All” tab of search results on the Twitter web client and Twitter for iOS & Twitter for Android apps. Over time, you’ll see more Tweets from this index appearing in the “Top” tab of search results and in new product experiences powered by this index. Try it out: you can search for the first Tweets about New Years between Dec. 30, 2006 and Jan. 2, 2007.
The full index is a major infrastructure investment and part of ongoing improvements to the search and discovery experience on Twitter. There is still more exciting work ahead, such as optimizations for smart caching. If this project sounds interesting to you, we could use your help – join the flock!
Acknowledgments
The full index project described in this post was led by Yi Zhuang and Paul Burstein. However, it builds on multiple years of related work. Many thanks to the team members that made this project possible.

5 No-Brainer Reasons Google Plus Is Right for Your Business by Janet Johnson

Many businesses are still skeptical about Google+.
They don’t know if having a Google+ business page will actually add any value.
I don’t blame them — Google hasn’t done the best job explaining the network.
Google’s social network is important — for many reasons!
And here are 5 no-brainer reasons to get your business on Google+ today.

5 No-Brainer Reasons Google Plus Is Right for Your Business

1. You’re On Google+ Already — and Don’t Even Know It

Do you have a Gmail account? Then you’re on Google+.
When you set up a Gmail account, you automatically set up a Google+ profile, too.
To get to Google+, log into your Gmail account. In the upper right corner you’ll see your name (or the name that you used to set up the account) with a + sign in front of it. Click it to be taken to your Google+ profile.
google-plus-business
It’s very important to at least add a good head shot & fill in your profile “About” info.
How should you fill out your profile? Here’s an excellent explanation:

2. Local Google Listings Connect to Google+

For a local business, a Google Places listing is one of the most important online listings you can have. Google gives preference to the Places listing when a city is entered into the search bar.
But, what many businesses don’t realize is that Google integrated the Places listings with Google+ a while back.
Now, when a Google Places listing is clicked it links to what’s called a “Google+ Local Page”. This page must be filled out by the business. Many haven’t been & some even link to the wrong Google Places page.
An easy way to find out which Google+ Local page links to your Google Places is to type the name of your business in the search engine. Find the Places listing (it should look like the below example) and click “Google+ Page.”
google-plus-business
If you’ve properly set up your Google+ page & linked it to your Places listing, it should contain a profile picture, cover photo, about information, links, etc.
Here’s an example of what not to do:
google-plus-business

3. It’s Easy to Connect with Potential Clients & Influencers

Other social networks have become so big & noisy that it’s hard to actually have meaningful conversations. That’s where Google+ excels.
People on Google+ take their time to converse.
For a major social network, Google+ is still small. It’s not as crowded as Facebook, which makes it a lot easier to chat with others.

4. It Helps Your SEO

The people you have in your Google+ circles & interaction you receive on Google+ posts can affect how you & your business appear in Google search results.
In other words, Google+ IS Google. Need I say more?!
Using Google+ can improve your SEO!
Your Google+ posts are indexed by Google & can appear in search results. And what business doesn’t want to be found on page one of Google?
Below is an example of a Google+ post that was indexed on page one of the search results.
google-plus-business

5. Hangouts On Air Are AWESOME!

If the first 4 reasons in this article don’t attract you to Google+, the Hangouts On Air feature should because it’s a game changer!
All businesses should use video to build rapport with clients & potential customers. But you used to need lots of equipment & it took lots of time & energy to create videos.
Not anymore!
Now you just need a smartphone & Hangouts On Air.
Google explained Hangouts On Air like this:
With Hangouts On Air, you can broadcast live discussions & performances to the world through your Google+ home page & YouTube channel. You can also edit & share a copy of the broadcast.
The possibilities are endless. Here are just a few of the powerful ways you can use Hangouts On Air:
  • Meeting / Consultation
  • Video
  • Interview
  • Client Testimonial
  • Show — Video or Podcast
Hangouts are automatically recorded right to your YouTube channel. When you’re done recording, you can edit right inside YouTube OR download, edit & upload.
Be sure to add keywords in the title, body & tags to give you a better chance of getting indexed by the search engine. Since Google has given preference to Hangouts On Air over other videos uploaded to YouTube, your chances of front page results are even greater with a Hangout!
I edited this Hangout with a few keywords & within minutes, these were the results:
google-plus-business

Conclusion

Unfortunately, bad press gave Google+ a bad rap — and undeservedly so!
Small business owners are afraid of Google+… and that fear must end. Google+ gets better every day.
If your business cares about how Google & video affect your results & sales, then I think it’s best to look deeper into what this platform can do for you.
Google+ has been a tremendous asset for me!

20 Ways B2B SEOs Can Leverage Schema.org Markup by Derek Edmond

knowledge-graph-brain-ss-1920
In a Searchmetrics report earlier this year, Schema.org – U.S. 2014: Rich snippets with HTML Microdata and RDF, it was revealed that while less than 1% of all domains researched had schema.org integrations present, almost 41% of keyword search queries contained a result snippet derived from schema.org markup.
Schema.org markups in Google SERPs for almost 40% of keywords investigated
More importantly, domains with schema.org integrations had significantly higher SEO visibility scores than those domains that did not include them.
Even with these findings, B2B marketers might remain skeptical, given that the most popular examples of schema.org integrations tend to be more consumer oriented. TV Reviews, Product Ratings, and Recipes are popular examples, but these obviously have a greater impact in the B2C space.
But, there is hope for B2B SEO, as well.
In this column, I will discuss 20 schema.org integration opportunities for B2B search marketers, organized by broader topics and how they can impact online marketing goals and objectives.
Note that this column reviews concepts and opportunities. If you’re looking for specific code references, BuiltVisible has a nice blog post detailing markup examples. Some of the links I provide to schema vocabularies also contain examples.

Quick Review: What Is Schema?

Schema.org is a tagging vocabulary that marketers and website owners can use to mark up their HTML pages in ways recognized by major search providers.
On-page markup enables search engines to understand the information on web pages and provide richer search results in order to make it easier for users to find relevant information on the web. Markup can also enable new tools and applications that make use of the structure.
What Google says about schema.org:
Search engines are using on-page markup in a variety of ways – for example, Google uses it to create rich snippets in search results…over time you can expect that more data will be used in more ways. In addition, since the markup is publicly accessible from your web pages, other organizations may find interesting new ways to make use of it as well.
Schema.org markup enables B2B marketers to better distinguish their web content in search engine results, with a goal of driving greater click-through rates and improving the exchange of information across search engine and social media platforms.

Web Page Elements

The first step in schema.org integration for B2B SEO is in establishing a foundation. The following vocabularies provide this support.
Schema Vocabularies: WebPage, Article, and Blog / BlogPosting
The WebPage vocabulary can be further broken down into more specific web page types like “AboutPage,” “ContactPage,” and “ProfilePage” in an effort to better differentiate web page objectives.
I prefer to place the schema code for “WebPage” in the opening tag and Article or Blog schema code literally around specific content on the web page. I’ve noticed that if you don’t use the WebPage markup but have other types of markup on the page, social platforms might confuse titles and descriptions when trying to share information.
Schema Vocabulary: Site Navigation
As covered at SMX this past Summer, Jeff Preston of Disney Interactive showed an example of where they used schema site navigation markup to influence site links in search results. This is a (relatively) simple tagging update that could produce positive gains in click-through rates if implemented in the navigational architecture.
Schema.org site navigation

Image & Video (Media) Objects

Schema Video Example
Schema Vocabularies: ImageObject and VideoObject
Searchmetrics’ 2014 SEO Factors report highlighted that web pages closer to the top of search results tended to contain a greater number of media assets, images in particular.
Searchmetrics SEO Factors: Media Usage
Per the Searchmetrics report:
Photos and videos not only make text more attractive for users, but for Google, this trend is likely to develop positively and be capped at a certain level.
The key takeaway for B2B SEO is that creating markup for images and video provides the opportunity for search engines to better understand the media assets found on the web page, which ultimately may provide a boost in organic search visibility.

B2B E-Commerce

We’re seeing more B2B organizations getting comfortable offering a mechanism for customers to purchase inventory online. As B2B marketers evaluate their options for online purchasing, the integration of schema markup is a factor not be overlooked.
Schema Vocabularies: Product, Offer, and Rating
Product and offer vocabularies provide the framework for a more comprehensive experience in search engine results specific to e-commerce catalog information. Ratings, coupled with the media-object elements listed above, provide an opportunity for B2B e-commerce vendors to stand out in search engine results, even if top placement isn’t possible.
Schema Ratings Example
Schema Vocabularies: Enumeration and SearchAction
Another way schema vocabularies can aid B2B marketers in e-commerce is in helping search engines identify functionality inherent to the shopping process. Enumeration could provide further clarity for a list of site-specific search engine results, or product inventory associated to a category or sub-category of the e-commerce catalog.
Implementation of the SearchAction vocabulary provides a method for site owners to ensure searchers land on their own sites and do not remain in Google search results if a sitelinks search box appears in Google results for branded queries.
This is specific to a recent Google announcement highlighting functionality that is designed to allow users to reach content on a third party site, directly through Google-generated site-search pages.
Google Site Search Example
AJ Kohn has a more in-depth blog post on potential issues with the sitelinks search box for large online brands.
Site owners should evaluate their own branded search results to consider the implementation of schema tagging for their site search functionality.
Even if your organization does not feature a sitelinks search box for branded queries, B2B marketers should consider integrating the SearchAction vocabulary (and how their internal search functionality performs) as a preventive measure down the road.

Localization

Another aspect of organic search engine results that B2B SEOs need to pay more attention to is local business listings. This is especially true for organizations working with multiple offices, dealer/distribution sites, and even franchise operations – all can benefit from more pronounced visibility for local search queries.
Schema Vocabularies: LocalBusiness and PostalAddress
I would recommend all organizational address information found on a website be integrated with schema LocalBusiness markup since it helps provide a more valuable search experience for brand-based queries, as well.
Schema Address Example
As detailed in an article written by Position2 this past summer, any event with a certain time and location can use the event markup vocabulary. Repeated events may also be structured as separate event objects.
Schema Event Markup
Schema Vocabulary: Event
This markup is ideal for the B2B organization hosting webinars, user conferences/roadshows, and any other events with specific time and location information. B2B marketers can join event markup with local address and/or offer markup to create an opportunity for even more comprehensive search listings, as well.

Brand Development

Brand awareness is a critical objective for most B2B marketers. The following schema vocabularies help marketers better identify their brand and key thought leaders in the organization.
Schema Vocabularies: Brand, Organization and Person
At the least, the schema vocabulary for Brand helps search engines distinguish an organizations’ logo, product information and align with a particular web address. The Organization schema can be used in association with local address information, on contact information or in other supporting marketing collateral.
Arguably, the Person schema can be used as a supporting mechanism for establishing the thought leadership of key organizational personnel. This could be particularly important now that Google no longer supports authorship in search engine results.
Through the Person vocabulary, B2B SEO’s can organizational leaders with information such as job title, images, contact information, and related web pages (such as social media profiles).

Email Marketing

According to resources from Google Webmaster Tools, the Structured Data Markup Helper can show marketers how to mark up emails that contain schema vocabularies, such as events (previously detailed), “go-to-actions,” and information highlights.
Gmail Markup Vocabularies: Go-To Actions, Order Information and Offers
These integrations allow B2B SEO’s and online marketers to create more comprehensive experiences in HTML-based email communication while providing Google with a greater understanding of context from the publisher. While we’ve just begun exploring this functionality, the opportunities look exciting.
Add Schema.org markup to emails
Go-To-Actions provide an opportunity to cross-link back to the publisher destination right in the preview pane of Gmail. Gmail grid view may display offer markup in an alternative, more visual way in the Promotions tab.
Order information markup is for exactly what it sounds like, allowing email recipients to receive more visual information designed to be easier for review.
Gmail Markup
Go-To-Actions and Offers provide for direct connectivity (linking) to a publisher’s applicable web address and Google+ page.
Order information ties back to the publisher’s order confirmation page and organizational information.
I suspect Google will also associate additional schema vocabularies present in HTML based email, as well (i.e., Brand, Organization, etc).
It should be noted that the schema vocabularies specific to email marketing and used by Google are still going through the standardization process of schema.org. They may change or cease to work in the long run.

Final Thoughts

The above are integrations of schema vocabularies we have implemented and/or begun discussions in integration with organizations we work with. I know there are many other opportunities for schema and great examples online.