Engineering Archives

Taboola Blog
Engineering

Javascript

20 October 2021

The Challenges of 3rd Party Scripting

Rbox, our recommendation product, is a 3rd party service embedded in publisher sites

Big Data

7 October 2021

Samplex: Scale Up Your Spark Jobs

In this article you will learn what Samplex is and how it is used to make processing of large raw datasets more efficient.

Building a POC for a Puppy: How Product Methodologies Help Guide My Decisions at Work and at Home

Culture

20 September 2021

Building a POC for a Puppy: How Product Methodologies Help Guide My Decisions at Work and at Home

Many R&D buzzwords and acronyms can seem like complex jargon — unnecessary shortcuts for concepts that are already pretty basic.

Work From Anywhere Support: The Story of Our IT & Support Work-From-Home Transition

Culture

31 August 2021

Work From Anywhere Support: The Story of Our IT & Support Work-From-Home Transition

During the pandemic, most companies quickly adapted and moved to a work-from-home model, as a sudden necessity of the lockdown restrictions introduced by efforts to combat the spread of COVID-19.

We Have Only Just Begun: A Message from Our SVP of Research & Development

Culture

16 August 2021

We Have Only Just Begun: A Message from Our SVP of Research & Development

We announced our plans to acquire Connexity, bringing eCommerce recommendations to the open web. And this is just the beginning.

Big Data

27 May 2021

ScORe – Schema On Read for Spark SQL

The world is not flat, it’s highly nested With over 4 billion page views per day and over 100TB of data collected daily, scale at Taboola is no joke. Our primary data pipe deals with masses of data and endless read paths. Could we optimize our schema for all these read paths? Guess not… Our schema is HUGE and highly nested. After digesting the data, we keep it in hourly Parquet files on HDFS, where each hour consists of about 1-1.5TB of compressed data. Our schema roughly looks like this: root |– userSession: struct | |– maskedIp: long | |– geo: struct | | |– country: string | | |– region: string | | |– city: string | |– pageViews: array | | |– element: struct | | | |– url: string | | | |– referrer: string | | | |– widgets: array | | | | |– element: […]

System

7 April 2021

TEST in PRODUCTION – should you?

You wrote your code. You even tested it. And now, you are eager to git push it. But how can you verify that it really works? In Taboola, we test our code in production! In this article, you will see how every software engineer, even on the first day in the company, can test in production – all thanks to a dedicated Jenkins pipeline job and lots of metrics. How hard is it to test in production? Quite hard. You probably already knew that. Everybody fears that moment when they need to test changes in production. The main reason is that not everyone has the required IT skills. Moreover, people have to repeat error-prone, manual tasks – which might result in downtime and revenue loss. For our release engineers, it was also an unmanageable headache – a “thundering herd” of developers eager to test their features in production. […]

The Challenges Of Uploading 150TB/day From Spark To BigQuery – Part 1

Big Data

25 March 2021

The Challenges Of Uploading 150TB/day From Spark To BigQuery – Part 1

Have you ever tried building an infrastructure to upload 150TB a day? Have you ever tried querying over 13PB without going bankrupt? These are some of Taboola’s PV2Google (pageviews to Google) service scale challenges that we deal with in our day to day. In this blog series, we’ll share how we do it, and the challenges we face. In this article (part 1) we’ll focus on the architecture. Part 2 covers the lessons we’ve learned over the years. Hello, Pageviews! Taboola’s goal is to power recommendations for publishers and advertisers. Our platform serves over 360 billion content recommendations and processes over two billion pageviews a day. Pageview is a record describing recommendations, user activity (such as a click), and much more on a user’s visit to a webpage. Currently, the pageview record has about 1,000 fields. Two billion pageviews generate a huge amount of data. This data is processed […]

The Challenges Of Uploading 150TB/day From Spark To BigQuery – Part 2

Big Data

25 March 2021

The Challenges Of Uploading 150TB/day From Spark To BigQuery – Part 2

In part 1 of the series we shared the architecture of Taboola’s PV2Google service which uploads over 150TB/day to BigQuery. In this article (part 2), we’ll share the challenges and lessons we’ve learned over the course of a few years. Lesson 1: queries might be (extremely) expensive We continuously upload pageviews to BigQuery and keep them for six months. This translates to over 13PB of pageviews in BigQuery. Querying the entire dataset would be extremely expensive, about $65K/query (assuming $5/TB). We apply a few methods and guidelines to substantially reduce this cost: Never use `SELECT *`: BigQuery’s query cost is based on the size of the data scanned. Most queries actually need only a few fields. Hence, selecting only the relevant fields will dramatically reduce the cost of the query. Cluster tables: clustering is a neat BigQuery feature that reduces the scanned row count. With clustering, BigQuery optimizes the data […]

System

3 November 2020

Don’t just run DNS, run it FAST

By Ariel Pisetzky and Tarek Shama Taken for Granted Taken for granted. That’s the way most users and even techies think of DNS. Or more precisely, they just don’t think of it at all. DNS is one of those things that for most users is a solved problem. You have a server with very reliable and stable software that can run for a very long time with little maintenance. The resolvers even have a nice built in failover mechanism for a secondary server. So, what more is there to say about this subject? Performance. Performance with DNS services has been seen as a geographical issue for years now. Yes, there have been paid DNS services that are faster at the DNS search level itself, especially if you are talking about complex records that have logic attached to them. Yet, the popular discussion is mostly around the global DNS providers. In […]

« First «...2 345 6...»Last »

Engineering

Start Your Taboola Career Today!