Victoria Seoane – Montevideo Labs https://www.montevideolabs.com Thu, 23 Nov 2023 19:03:58 +0000 en-US hourly 1 https://wordpress.org/?v=7.0.4 https://www.montevideolabs.com/wp-content/uploads/2023/07/cropped-iso-Mlabs-gris-32x32.png Victoria Seoane – Montevideo Labs https://www.montevideolabs.com 32 32 Some Tips For re:Invent 2023 https://www.montevideolabs.com/2023/11/22/some-tips-for-reinvent-2023/ Wed, 22 Nov 2023 17:22:34 +0000 https://www.montevideolabs.com/?p=10786 Some tips for re:Invent 2023 (that I wish I had known in previous years!)

AWS re:Invent is just around the corner! With an expected attendance of 50,000 participants, this year’s conference promises to infuse the city with a wave of innovation, fresh ideas, and a vibrant sense of community. As you prepare to embark on this exciting journey, we’ve put together a curated list of recommendations to enhance your experience and ensure you extract maximum value from this extraordinary event. So, as you pack for this grand adventure, let our tips guide you in making the most of every moment at AWS re:Invent!
 

Failing to plan is planning to fail

With such a wide variety of sessions, it’s understandable to want to go to everything. However, it’s crucial to review the session catalog beforehand. I’d recommend listing all the sessions you definitely don’t want to miss, and then listing “nice-to-haves”. Once you have your list of musts, try to find reserved seating for those. If you can’t find reserved seating, you can join the queue outside the session – you’ll usually be able to get in.

Also, with 6 venues across the Las Vegas Strip, planning is essential. My first time around, I thought that hotels were close by, and that I could just “cross the street” between sessions. Big. Mistake. The venues are huge, and some talks can be in different areas of the same hotel. My recommendation here would be to try and choose one hotel for the day, or at least one hotel for the morning, one for the afternoon.

Now – how to choose? It’s important to keep in mind that all sessions have “levels”: from 100 to 400. 100-level sessions are usually basic, high level overviews – a great opportunity to learn about something new. 300 and 400 sessions require that you have some previous experience on the subject matter, but in my experience they have been the most thought-provoking. They’re also the sessions where you will find people with similar interests.

There are different types of sessions: Technical Sessions, Hands-on workshops, Chalk-talks, builders sessions and keynotes addresses. I’ve found that Chalk talks are usually engaging and lead to a lot of interaction between speakers and the audience. The keynotes are usually fun, and they usually announce new services. Werner Vogel’s keynote is usually my favorite – and last year’s Matrix analogy didn’t disappoint!

Something I learned last year – remember that new sessions will pop up, even during the conference. It’s important to review your plans each night – and keep a flexible attitude 🙂 The AWS Events app is a great tool – make it your best friend. If you’re in a venue, and you couldn’t go into a session you were looking forward to, you can easily find sessions close to you simply filtering by venue / interests / day.

There is such a thing as too much of a good thing

One of the mistakes I made last time was to plan for a jam-packed schedule. With such an amazing variety of sessions, I simply couldn’t pick a few. The first day, I went to three sessions in the morning, four sessions in the afternoon – and by 8 pm, I was exhausted. I realized that there is a limit to how many sessions you can pay attention to in a day. My rule of thumb would be to try to plan four or five talks per day – and try to enjoy the conference. For instance, go to the expo to check out the amazing booths, or grab lunch at the big dining rooms instead of grabbing a packed lunch. Recharging is crucial to make the best of the sessions, engage in Q&A, and join the night events, which are always fun!
 

Befriend the shuttle

I love to walk outside; I find it clears my mind and improves my mood. However, I’ve found that you do a ton of walking inside the venues, so it’s a good idea to use the shuttle as much as possible. The shuttle is a conference service that connects each venue to the rest. It has the advantage that it drops you off at the closest door to the conference. If you opt to walk around Las Vegas, you might find yourself walking ’ longer than necessary. I’ve chosen to do it on occasion – but it can be a waste of time.

Also, you don’t need to be staying at a conference hotel to use the shuttle early in the morning. With your attendee badge, you can join the shuttle at the closest stop.

Speakers love feedback and questions

Don’t be afraid to reach out to the speakers at the end – even if it’s just to congratulate them. In my first re:Invent, I was hesitant to stay back and talk with the speakers. However, on my second conference, I made a point of sticking around, reaching out to the speakers, asking follow up questions. It was great fun, I met amazing people and built a great network in the process.

Along this same line – ask questions at the sessions – you can start great discussions, and it’s a way to show speakers that you enjoyed the session and were actually listening. It’s terrible to finish speaking, ask if somebody has a question, only to hear silence. Speakers probably wonder if anyone enjoyed the talk, or if it sparked new ideas in their audience. Asking questions is a way to delve deeper into topics you want to learn more about.

Lunch tables, queues and the shuttle

What do these have in common? Opportunities to meet new people! Networking is a significant aspect of re:Invent, and you never know where you might find your next collaborator, mentor, or friend. Don’t hesitate to strike up conversations with fellow attendees, whether you’re waiting in line for a session, riding the shuttle, or having lunch. These impromptu interactions can lead to valuable insights, shared experiences, and even potential partnerships. Seize every opportunity to connect!

Go into conference mode – Mute your notifications, charge your devices – and mute your phone during sessions.

One of the keys to fully immersing yourself in re:Invent is to mute your notifications and silence your phone during sessions. The sheer volume of information and networking opportunities can be overwhelming, and constant digital interruptions might hinder your ability to focus. By muting notifications, you can ensure that you’re fully present in the moment, absorbing the valuable insights and making meaningful connections. Also, you don’t want to be that person that interrupts speakers with a phone call, do you?

Another thing to consider: hotels are usually short on power outlets! Remember to charge your devices overnight, or you’ll find yourself walking up and down aisles, and sitting on the floor next to the only outlet you can find. On the other hand – it is a great conversation starter!

An unexpected essential: chap stick

Las Vegas, the host city for re:Invent, is known for its dry climate. Combine that with spending long hours indoors in air-conditioned environments, and you’ll quickly understand the importance of keeping your lips moisturized. Pack plenty of chapstick to ensure that you’re comfortable throughout the event. Small details like this can make a significant difference in your overall experience.

The best adventures are usually beyond your comfort zone

While it’s natural to attend sessions directly related to your expertise, don’t shy away from exploring talks outside your comfort zone. re:Invent offers a rich variety of sessions covering a broad spectrum of topics. If you’re an absolute rookie in a particular subject, 100-level talks are perfect for you. Venture into areas that might be unfamiliar but intriguing. You’ll not only broaden your knowledge but also gain fresh perspectives that could inspire innovative solutions in your own work.

Re:Invent is an unparalleled experience, intertwining learning, networking, and sheer enjoyment. Reflecting on the inaugural night of re:Invent 2022, I found myself gathered with a friend, furiously scribbling down notes—our minds brimming with newfound knowledge to the point where our handwritten scrawls bordered on illegible, a testament to the wealth of insights gained. Little did I anticipate the exhilarating slide down a DataDog slide, the impromptu camp set up on the Venetian floor — a vivid illustration of the unexpected and delightful nature that defines the essence of re:Invent.

Stay ahead of the curve on the latest trends and insights in big data, machine learning and artificial intelligence. Don’t miss out and subscribe to our newsletter!

]]>
Detecting Brand-Unsafe Content Through Computer Vision https://www.montevideolabs.com/2023/10/20/detecting-brand-unsafe-content-through-computer-vision/ Fri, 20 Oct 2023 13:39:44 +0000 https://www.montevideolabs.com/?p=10576 In the realm of digital advertising, ensuring brand safety is a paramount concern. It involves the task of identifying and safeguarding ad placements within content that aligns with a brand’s values and objectives and – sometimes more importantly – detecting content that doesn’t align with a brand’s message. In 2019, the Global Alliance for Responsible Media (GARM), released the brand safety and suitability standard, which helps brands communicate their particular needs in a shared language across the industry.

Business challenge

With this context in mind, we want to present a challenge that was brought to us by one of our clients: how to identify brand unsafe content in a large corpus of videos in a cost effective way? Video analysis was cost prohibitive, so we designed a solution that leveraged image and audio analysis. After all – videos are a large set of images overlaid with sound, right?

In this article, we’ll focus on the computer vision component of the analysis – how to find, among other things, video scenes related to crime, weapons, terrorism, alcohol, drugs or pornography?

First step – designing the solution

The first thing we decided was to base our solution upon a sample of video stills from the entire video. Even though this increased the chance of error, image analysis is less expensive than video analysis, so we could sample a larger number of images. For illustrative purposes, if we used AWS Rekognition for content moderation using video analytics it costs $0.10/min, while image recognition would cost $0.0008 per image. We can sample 100 images in a minute of video, and still be below video costs. The number of images to sample will be a parameter to consider, in order to balance model accuracy and overall cost.

To sample these images, we wrote an Open CV script to load a video and sample a subset of stills from it.

import cv2import os
def sample_images(path, file_name, directory_to_store_at=”/tmp”): “”” Capture frames from a video and save sampled images at specified intervals.
Args: path (str): Path to the input video file. file_name (str): Base name for the output image files. directory_to_store_at (str, optional): Directory to save the sampled images. Defaults to “/tmp”.
Returns: List[str]: List of file paths to the sampled images. “”” images = [] cam = cv2.VideoCapture(path) total_frames = int(cam.get(cv2.CAP_PROP_FRAME_COUNT)) fps = int(cam.get(cv2.CAP_PROP_FPS)) video_duration = total_frames / fps
# Calculate the interval (in minutes) to sample at # according to the video duration (adjust to business logic) interval_minutes = 0.1 if video_duration < 60: interval_minutes = 0.1 else: interval_minutes = 1 frame_interval = int(fps * 60 * interval_minutes)
if not os.path.exists(directory_to_store_at): os.makedirs(directory_to_store_at)
# start sampling at frame 15 to avoid the initial black frames currentframe = 15 while(True): current_frame = current_frame + frame_interval cam.set(cv2.CAP_PROP_POS_FRAMES, current_frame) ret, frame = cam.read() # reading from frame if ret: # if video is still left continue creating images file_id = file_name + ‘frame’ + str(current_frame) + ‘.jpg’ name = directory_to_store_at + file_id image_string = cv2.imencode(‘.jpg’, frame)[1] with open(name, ‘wb’) as file: file.write(image_string.tobytes()) images.append(name) else: break # Release all space and windows once done cam.release() cv2.destroyAllWindows() return images

Once we have our images sourced from stills, how do we want to classify them? What model should we use? Our first thought was to leverage a similar idea to the one we exposed in our previous blog post, and classify them into all the different categories, and have a “catch-all” category for safe content. However, this had several problems. For starters, an image could belong to several different categories – making multi-class classification an insufficient solution. Also, what would happen with images that belonged to none of the categories? The “safe” category would potentially be huge. The amount of data we would need to train the models on all the things that are actually safe would’ve been massive. Not only that, we would potentially be incurring in class imbalance and the problems that this can entail. For example, the model may become biased towards the majority class, leading to poor performance on the minority class.

To solve the multiple-label issue, we considered multi-labeled models, but the issue with the number of examples needed for “safe” content still remained.

Because of this, we decided to tackle the problem from a new perspective. Instead of building one model to classify them all (sorry, easy joke), we built a set of one-class classification models, one for each of the classes we want to detect. One-class classification, also known as anomaly detection, is a technique where the model learns to identify a specific class of data from a predominantly normal dataset. Autoencoders, a type of neural network architecture, are particularly well-suited for this task.

Let’s take a look into Autoencoders!

 
Autoencoders are a type of artificial neural network used in unsupervised machine learning. They are primarily designed for dimensionality reduction, feature learning, and data compression. Autoencoders have applications in a variety of domains, including image and signal processing, anomaly detection, and even natural language processing.

An autoencoder consists of two main parts: an encoder and a decoder. Here’s how they work:

Encoder: The encoder takes an input (which can be an image, a sequence of data, or any other form of structured data) and maps it to a lower-dimensional representation, often referred to as a “latent space” or “encoding”. This process involves a series of transformations and layers that capture essential features or patterns in the input data. The encoder’s goal is to reduce the dimensionality while preserving the most important information.

Latent Space: The latent space is a compressed representation of the input data, and it typically has a lower dimension than the original data. It serves as a bottleneck in the autoencoder architecture, forcing the model to capture the most critical features in a compact form.

Decoder: The decoder takes the latent space representation and reconstructs the input data from it. It consists of layers that perform the reverse operation of the encoder, attempting to generate an output that closely resembles the original input. The quality of the reconstruction is a measure of how well the autoencoder has learned to capture the essential features of the data.

Autoencoders are trained using a reconstruction loss, which measures the difference between the input and the output. The objective during training is to minimize this loss, effectively teaching the autoencoder to reproduce the input data as accurately as possible. In the case of one-class classification, if the reconstruction loss is high for a particular input, it suggests that the input does not fit well with the learned patterns, indicating that the image doesn’t belong to the class that the model was trained on. Let’s use an example: Imagine we train two auto encoders, one for detecting weapons and another one for detecting drugs. We can then submit an arbitrary image to each auto encoder and consider the reconstruction error for both. If we find the auto-encoder for drugs has a low reconstruction error, then it is likely that the image is related to drugs.

Sourcing the necessary training data

It’s hard to find open source datasets on many of these topics – especially due to their sensitivity. This is why it was crucial to create our own datasets. We created eight different datasets – crime, arms, military scenes, terrorism, explicit content, sexy content, smoking images and drinking images; each one having between 200 and 500 images. To do this, we used commercial-use allowed images from varied sources, such as Flickr, Google and Wikimedia. These were subject to a manual review from our team, to ensure that the source images were good representatives of what a brand unsafe video could have. For instance, an image of a soldier posing in uniform is related to military topics, but is not what the GARM standard refers to as unsafe content.    

A personal aside – the data labeling step is a task that is best to do as a team and in small doses, and if possible, looking at videos of adorable puppies in between.

Preparing the data for training

We’ll place all of our images in a folder, defined here as image folder.
We’ll also define an image height and width, which is necessary for the autoencoder models. It’s important to use an image size that keeps the image quality without it being too large – since large images will create larger models – and require more memory to run! In our case, 256×256 pixels is a reasonable shape.
import osimport numpy as np
# Set the path to the folder containing your imagesimage_folder = ‘./training_images_military’
# Define image dimensionstarget_height = 256target_width = 256
# Initialize an empty list to store the preprocessed imagespreprocessed_images = []from keras.layers import Input, Densefrom keras.models import Modelfrom keras.preprocessing.image import load_img, img_to_array
# Loop through the images in the folderfor image_filename in os.listdir(image_folder): # Construct the full path to the image file image_path = os.path.join(image_folder, image_filename) # Load the image using Keras’ load_img function image = load_img(image_path, target_size=(target_height, target_width)) # Convert the image to a NumPy array and normalize pixel values image_array = img_to_array(image) / 255.0 # Append the preprocessed image to the list preprocessed_images.append(image_array)
# Convert the list of preprocessed images to a NumPy arrayX_train = np.array(preprocessed_images)

Model training

We’ll now create a function to train our autoencoder.

In this case, we chose to use Dense layers, but there are different layers available: among others Conv2D, MaxPooling2D, UpSampling2D. The architecture chosen was a combination of study and experimentation – figuring out which architecture yielded the best results while keeping the model complexity reasonably low. The optimal batch size and epochs were a combination of high accuracy and overall model size – we wanted good models but that could run on computers with 256GB of memory.
def train_and_evaluate(epochs, batch_size, loss, training_data): “”” Train and evaluate an autoencoder model on the provided training data.
Args: epochs (int): The number of training epochs. batch_size (int): The batch size for training. loss (str): The loss function to be used for training. training_data (numpy.ndarray): The training data to be used for autoencoder training. target_height (int): The height of the target input images. target_width (int): The width of the target input images.
Returns: tensorflow.keras.models.Model: Trained autoencoder model. “””
input_dim = (target_height, target_width, 3) dim = target_height * target_width * 3 input_dim = (dim,) input_img = Input(shape=input_dim)
encoded = Dense(units=256, activation=’relu’)(input_img) encoded = Dense(units=128, activation=’relu’)(encoded) encoded = Dense(units=64, activation=’relu’)(encoded) decoded = Dense(units=128, activation=’relu’)(encoded) decoded = Dense(units=256, activation=’relu’)(decoded) decoded = Dense(units=dim, activation=’sigmoid’)(decoded) autoencoder = Model(input_img, decoded) autoencoder.compile(optimizer=’adam’, loss=loss)
# Train the autoencoder epochs = epochs batch_size = batch_size validation_split = 0.2 history = autoencoder.fit(training_data, training_data, epochs=epochs, batch_size=batch_size, validation_split=validation_split, verbose=0) return autoencoder
With that function defined, we trained the model:
autoencoder = train_and_evaluate(50, 16, ‘mse’, X_train)

Testing the model

Now, here comes the fun part. We’ll use the following function to classify our images!
def classify_image(image_path, autoencoder, threshold=0.001, target_height = 256, target_width = 256): “”” Classify an image based on its reconstruction error using an autoencoder.
Args: image_path (str): Path to the image file to be classified. autoencoder (tensorflow.keras.models.Model): Trained autoencoder model used for image reconstruction. threshold (float, optional): Threshold for classification. Images with a Mean Squared Error (MSE) below this threshold are considered to belong to the class. Defaults to 0.001. target_height (int, optional): The target height for resizing the image. Defaults to 256. target_width (int, optional): The target width for resizing the image. Defaults to 256.
Returns: tuple: A tuple containing: – mse (float): The Mean Squared Error (MSE) between the original and reconstructed image. – belongs (bool): True if the image belongs to the class based on the MSE, False otherwise. “”” image = load_img(image_path, target_size=(target_height, target_width)) image_array = img_to_array(image) / 255.0
# Reshape the image array to match the input shape of the autoencoder image_array = image_array.reshape(1, -1) # Use the autoencoder to predict the reconstructed image reconstructed_image = autoencoder.predict(image_array)
# Calculate the Mean Squared Error (MSE) as the reconstruction error mse = np.mean(np.square(image_array – reconstructed_image))
if mse < threshold: #Image belongs to the class belongs = True else: #Image does not belong to the class belongs = False return (mse, belongs)
And now, for the moment of truth – we test it on our data:
for image in os.listdir(image_folder): mse, belongs = classify_image(“training_images_crime/” + image, autoencoder, threshold)

Tuning the threshold

If you are paying attention, you’ll see a “threshold” mentioned in the code above, very casually. What is that? Remember some sections ago, we talked about how autoencoders can be used for one class classification? We ask the model to encode an image, and then measure the reconstruction error. In this case, we’ll be using Mean Squared Error (MSE) – the mean of squares of errors between real image and predicted image. If the error is large, it means that the image doesn’t belong to the set of images we used for training because the model wasn’t able to recreate it accurately. This is the concept we’ll use to classify our images. However – what is a “large error”? This is something we’ll need to investigate – set a threshold that is consistent with your data. An error that is low enough to keep True Positives and True Negatives high, and False Positive and False Negatives low. Of course, this will also depend on your business use case, and the sensitivity to Type I and Type II errors.

This is the distribution of the MSEs for the images that belonged to the class:

And this is the distribution of the MSEs for the images that did not belong to the class:

As we can see, the choice of threshold is both a science and an art – choosing a lower threshold will make us mark a lot of images that belong to the class as not belonging, and risk not catching brand unsafe content. On the other hand, choosing a higher threshold will make us mark more inventory than strictly necessary. A threshold of 0.05 in this case was a good choice, capturing most of the images that did belong to the class, while misclassifying very few images that didn’t belong.

An interesting exercise is to analyze the images that were misclassified. In this case, since many of the training images had a desertic background due to the military campaigns in the Middle East, images of the desert were classified as having military content. To improve this, I improved my training data, to capture images from other contexts, such as the Vietnam war.

Conclusion

Brand safety in advertising demands meticulous attention to the context in which ads are displayed. 
Our challenge was to identify how to detect potential ad placement context in a cost-effective manner.

We can address this by employing Keras autoencoders for one-class classification of video frames. This approach significantly reduces data labeling requirements, since multi-label and multi-class models would require a wide variety of “brand safe” data.

Training an autoencoder on brand-safe content and establishing a threshold for reconstruction errors allows us to accurately identify placements that conform to brand values.

This solution offers significant potential for streamlining brand safety efforts in advertising, ensuring that ads are presented in alignment with a brand’s identity and mission.

Stay ahead of the curve on the latest trends and insights in big data, machine learning and artificial intelligence. Don’t miss out and subscribe to our newsletter!

]]>
5 inspiring books for new managers https://www.montevideolabs.com/2023/07/04/5-inspiring-books-for-new-managers/ Tue, 04 Jul 2023 12:09:43 +0000 https://www.montevideolabs.com/?p=7195 This is an extraordinary era of technology, where algorithms, AI, and UX have become integral parts of every business and industry, bringing opportunities, innovation, and growth. All of these make us feel challenged, energized and highly motivated to get out of bed every morning and tackle new challenges. Companies grow, opportunities for leadership arise, and we start to question ourselves: am I ready for leadership? I studied to be an engineer – mostly “hard skills”. How can I prepare myself to lead others?

In my case, I’m an avid reader, so it was only natural for me to look for inspiration in books. Our office library is full of treasures, and I usually come back from trips with an extra heavy backpack, filled with books. So, without further ado, my top 5 recommendations!

Trillion Dollar Coach

This was the first book I read when I became VP of Engineering. I was going to be in charge of leading our internal coaching programme, and I must admit I felt horribly unprepared. Máximo, our CEO, recommended this book as a summer read, especially because of the great anecdotes it has. It tells the story of Bill Campbell, Silicon Valley’s most famous basketball coach turned business coach. Among his coachees we can find Steve Jobs, Larry Page, Sergei Brin, Mark Zuckerberg, Sheryl Sandberg, Tim Cook, Jeff Bezos, Mary Meeker, John Doerr, Ruth Porat, Scott Cook, Brad Smith, Ben Horowitz, Marc Andreessen… and the list goes on. (I know – my definition of summer read is questionable at best).

My favorite quote from the book is the following: “It’s the people. People are the foundation of any company’s success. The primary job of each manager is to help people be more effective in their job and to grow and develop.” I love it because it captures the essence of my job description with the right focus: People. It also gives us a guideline for making hard decisions: we should have our “first principles” – immutable truths that are our company’s foundation, and guide decisions from there. 

Something that is quite interesting about this book is that it is a collection of stories from his life, and the people whose lives were changed with his coaching philosophy. His biggest tenets were empathy, trust, collaboration, openness and continuous learning. The role of feedback and 1-on-1s is brought front and center, and many of the guidelines provided now work as starting points for my 1-on-1 meetings. 

Another concept that I liked was the difference between a coach and a mentor: mentors share their wisdom with us, while a coach pulls up their sleeves and helps us identify areas we can improve and hold us accountable for those improvements. By doing so, coaches help people realize their full potential, oftentimes becoming key players in articulating different styles, personalities and backgrounds.

As a final reflection, the most significant lesson and the one that resonates with me the most is that he “gave us permission to care”. Everyone he coached felt loved, genuinely cared for, and were both cheered on and challenged. This is also very linked to other books in the list, but it’s important to remember that: As managers, you will develop personal relationships with your team, and that is a good thing!

What you do is who you are

As a company, we are experiencing exponential growth. Taking care of our culture and its intentional development had a profound impact on my year. I found many gems in this book that presents anecdotes from a wide variety of places, such as the Haitian slave rebellion led by Toussaint Louverture or Genghis Khan’s leadership style.

I have to admit – I have always felt very proud of our culture as a team, but hadn’t reflected much as to what a strategic advantage it gave us as a company. The author emphasizes that without a strong company culture, it’s very hard to succeed in the modern business world. In our day and age, where decisions have to be made fast, a strong culture can work as guiding principles for our team. 

Setting clear “cultural norms and operating principles” help create consistency and alignment, and avoid headaches.

One aspect I personally liked, that set apart this book from many others of this same topic, were the anecdotes and histories mixed in with business recommendations, and how he dared to take culture out of the workspace and look for examples in other organizations. A prison gang leader might not sound like the best role model, but some of his decisions and mistakes can teach us a lot.

The author argues that our company’s culture is defined by the practices and behaviors – not just the values and beliefs. When leaders succeed in creating strong and distinctive cultures, these can work as a set of guidelines for everyone to follow – whether you meant it or not! He is not afraid of putting leaders on the hot seat, and stating plainly that if leaders don’t walk the talk and lead by example, words are useless. This is also shown in some of the historical anecdotes (but I won’t give spoilers!)

Finally – the role of rituals was something I particularly enjoyed. Back in the day in our office, on your first day of work, you had to assemble your chair with your teammates. This was a great ice breaker, and it led to lasting friendships (and some wonky office chairs). With our growth and the COVID pandemic, it’s something we stopped doing, but new traditions have started (such as music nights and soccer Fridays). It’s pretty cool to get reinforcement on things we love.

Radical Candor

Máximo, our CEO, mentioned this book in each and every one of our 1-on-1 meetings. Each time I had to prepare for a difficult conversation, he’d quote Kim Scott’s wisdom to help me prepare for it. Since he is also quite a book worm, the fact that this book had been so meaningful for him made me curious. So, I decided to give it a chance.

In a nutshell, this is a book about being a great boss – helping people grow while being a decent person. It focuses on honest and straightforward feedback and fluid communication – combining empathy, compassion and a will to challenge. I’ll try to summarize it:

As managers, we have two dimensions we’ll focus on for this framework:

  • Care personally
    • People come to work, and we develop a relationship with them
  • Challenge directly
    • We feel empowered to tell someone when their work is good – but also when it is not.

Let’s put this into an example: Someone comes to the office, and they have their shirt on backwards.

Don’t Care Personally

  • Challenge Directly: You laugh, point at them and say loudly: “Hey, did you get dressed in the dark?”
  • Don’t Challenge Directly: You whisper to the rest: “Can you believe that? If it happened to me, I’d realize it immediately”

Care Personally

  • Don’t Challenge Directly: You don’t do anything so as not to make them feel self conscious – it’s not a big deal, they’ll realize it eventually
  • Challenge Directly: You kindly take them somewhere quiet and tell them: “Hey, you’ve got your shirt on backwards – just wanted to let you know in case you want to deal with it”

What the book aims to highlight is that often, Radical Candor can make us uncomfortable. But when we care about someone, it’s important to challenge them directly to help them grow – else, how will they know they have a wardrobe malfunction?! And even if hard conversations are hard for us – they’ll be even harder coming from someone who doesn’t care as much.

Another point we want to make; caring about our teammates isn’t enough if we don’t dare to challenge them directly. If we really want to help them, it’s important to both care and challenge. 

You might think that some of the ideas aren’t groundbreaking – but the examples she gives (and she isn’t afraid of showing her own mistakes) helps drive the point home. She even tells an anecdote in which Sheryl Sandberg told her she was sounding stupid – not something anyone would like to admit in a popular management book, where your own voice and authority are key to your credibility (and the book’s!). 

Right now, I am reading her second book, Just Work, and it’s inspiring to see how she challenges some of her own ideas. I think that is a very brave thing to do – especially after penning a best seller. She mentions how it can be hard for women or minorities to implement Radical Candor without being labeled, and how many of her ideas and experience came from a place of privilege. Interesting read – and in her amusing, no B-S style, which is always refreshing!

When they win you win

After reading “Radical Candor,” I discovered a book written by one of the co-founders of Candor Inc., Russ Laraway. I liked the premise on the cover: “Being a great manager is simpler than you think.” I began reading it on a train from Boston to NY, and as snow fell outside my window, I was hooked. There was something about this book that made me feel like the author truly understood his craft. It felt as if I were having a conversation with an older sibling who had been through it all.
In the introduction, it talks about the three key elements of leadership: Direction, Coaching and Career. In his own words.

  1. Direction — Good managers ensure that every member of their team understands exactly what is expected and when it is expected.
  2. Coaching — Good managers coach their people toward both short- and long-term success, helping them understand what they should continue to do and where and how they can improve.
  3. Career — Good managers invest in their people’s careers in a way that considers their long-term goals and aspirations, beyond the four walls of the current company, and certainly beyond their next promotion.
  4. The premise is that as managers, our goal is to have engaged employees – and the manager seems to be the biggest driving force behind good (or bad) employee engagement. The author argues that if we focus on these three things, these will help our teams become engaged, and engaged employees deliver the expected results. This is shown throughout the books as 3 <-> E <-> R.
  5. One of the things I liked the most about this book was that it was eminently practical: It provides sample questions that we can use for self-assessment or to ask our teams to assess us. It also offers guidelines to help organize our 1x1s, (which was tremendously helpful when preparing our internal team’s training material.). Additionally, the book frequently references other books, demonstrating the author’s passion for team building and constant pursuit of learning.

Another aspect that resonated with me was when he said that being busy often competes with being focused. He emphasizes the need to learn how to subtract and to help our direct reports in doing the same. There’s an overall feeling that you must be busy to be productive, and it felt refreshing to read that it’s okay to do less and give permissions to do less. It usually leads to more productivity.  

I also liked the overall style. One of the first chapters has this hilarious title: “You just can’t suck as a manager.” Okay, please, tell me more! Another quote that I highlighted was: “First, know what you expect. Then, tell your team what you expect”. How many times have we failed for not being 100% aware of what was expected of us? Do we always know our success criteria?

The two last things I want to highlight are the following: he encourages us to use praise as a learning tool, and to always ask permission when we need to have a tough conversation. I liked them because they both prioritize our direct reports as human beings. It’s unrealistic to expect that everyone will always be in the best space to listen to criticism. Similarly, it’s a good idea to give praise when it’s due – positive feedback encourages us to continue doing what we were doing right. My personal style is very caring and hands on – so reading this helped me feel like I’m on the right track (at least in some aspects! 😀 )

Trust and Inspire

Have I mentioned I usually bring home three or four books from every trip? This year was no different. While trying to defrost on a very cold Massachusetts morning last December, we entered the Harvard Book Store with my husband.
Forty-five minutes later, we emerged with three (quite heavy) books, to the bemused looks of our parents. One of those books was Stephen Covey’s “Trust & Inspire”. I liked the premise of trying to “unleash greatness in others” – the idea that as a leader you empower others to grow, to achieve new heights.
This book works around a main idea: the world has changed, the nature of work, the workplace, the workforce and choice – all have changed.

In response to this new paradigm, the author states that leaders should focus on three main principles: who you are, how you lead, and how you help people be inspired. We are in an era in which teams need people who have more than just knowledge – we need knowledge, inspiration and passion. 


In this new era, Covey claims, the old “Command and Control” leadership style is out of date and out of place. We need a new style, which he names “Trust and Inspire”. To achieve this style, we need to model the behavior (again – very similar ideas we’ve heard before), you need to trust your team, and you need to help them find inspiration. A quote from Eleanor Roosevelt featured in the book summarizes this: “A good leaders inspires people to have confidence in the leader; a great leader inspires people to have confidence in themselves”.  
One aspect I really liked about this book was that it included “exercises” – for example, at the end of each chapter, it included notes that are easy to remember. For example: “There is enough for everyone, so my job as a leader is to inspire, not merely motivate”. These are great to keep at hand, or to use as training materials for new generations of leaders. 

Another interesting “gift” was that the author went one extra mile, and included an entire section on “barriers” we might see as we try to implement this. For instance, some reasonable doubts such as “these won’t work here”, or some fears such as “what if it doesn’t work” or “what if I lose control”. These fears are reasonable, and Covey includes a set of self-assessments to identify them, and ways to tackle them in a respectful way. 

Conclusion

So, that’s my list: the five books I found most inspiring for this journey I’m on. It’s unrealistic to expect that books alone will do the trick – now comes that fun part that is to select which aspects of all of these toolboxes we want to start working on, which ones apply best to your situation, and get moving! One of the best takeaways is that nobody is perfect – and even these great managers have enough mistakes to fill the pages of a book! So, let’s roll up our sleeves, and get to work!

Stay ahead of the curve on the latest trends and insights in big data, machine learning and artificial intelligence. Don’t miss out and subscribe to our newsletter!

]]>
Orchestrating an ML Workflow with Step Functions and EMR https://www.montevideolabs.com/2023/06/20/orchestrating-an-ml-workflow-with-step-functions-and-emr-2/ Tue, 20 Jun 2023 13:47:17 +0000 https://www.montevideolabs.com/?p=7016 As we have seen on previous posts (Launching an EMR cluster using Lambda functions to run PySpark scripts, part 1 and 2), EMR is a very popular big data platform, designed to simplify running Big Data frameworks such as Hadoop and Spark. Taking advantage of this platform allows us to scale up and down our compute capacity, leverage cheap spot instances and integrate seamlessly with our Data Lake. When we combine EMR with Lambdas, we can launch our clusters programmatically using serverless FaaS, and react to new data becoming available, or triggering on a cron schedule.

In this post, we propose to take this approach of using Lambdas to trigger EMR clusters one step further, and use AWS Step Functions to orchestrate our ML process. AWS Step Functions is a fully-managed service that makes it easy to build and run multi-step workflows. It allows us to define a series of steps that make up our workflow, and then Step Functions executes them in order. These steps can be run sequentially, in parallel, wait for previous stages to finish, or run conditionally if other workflows fail. The image below shows an example of a Step Function. It is a series of steps, where we can see a Start, Steps, Decisions, Alternatives, and an End. This is ideal for real-life workflows, where different decisions need to be made as the process progresses.

  • An interesting advantage of using Step Functions for Machine Learning workflows (instead of a single workflow, for example) is that we can spin up independent clusters for each step. Different stages of a Machine Learning pipeline can have different memory requirements and processing power. Being able to easily spin up a cluster for each step allows us to use the right size of instances for each  stage of the workflow, thus being more efficient with our time and spend.

Some other advantages to this approach are the seamless integration of StepFunctions with other services, such as Lambda, that allow us to build a complex workflow without having to orchestrate it manually, with easy out-of-the-box monitoring and error handling that StepFunctions provide – simplifying our retry logic and overall process monitoring.

So, hoping I have convinced you to try this new approach, let’s get hands on!

Spinning up our Step Function

The generation of the Step Functions can be done in two different ways: using the Step Functions visual editor, or leveraging CloudFormation. When I began this post, I wanted to use CloudFormation, but I ran into a CloudFormation limit (CloudFormation doesn’t allow us to create StepFunctions with a definition as long as I needed it to). This is why I will use the Step Functions visual editor, combined with Amazon States Language.

From the AWS website:

The Amazon States Language is a JSON-based, structured language used to define your state machine, a collection of states, that can do work (Task states), determine which states to transition to next (Choice states), stop an execution with an error (Fail states), and so on.

 

Our ML Workflow

We will define a workflow made up of 3 steps: data preprocessing, model training and model evaluation.

A set of examples, coded in Python using PySpark, and step by step to launch our first Step Function!

For this, we will need to create five things:

  • An S3 bucket where you will store your scripts: preprocessing, training and evaluation, which contain the logic for each of those steps:

preprocessing_script.py

import argparse
# Import necessary libraries
from pyspark.sql import SparkSession
from pyspark.ml.feature import VectorAssembler
from pyspark.ml.classification import RandomForestClassifier
from pyspark.ml.evaluation import MulticlassClassificationEvaluator
if __name__ == "__main__":
    # Initialize Spark session
    spark = SparkSession.builder.appName("PreprocessingApplication").getOrCreate()
    parser  =  argparse.ArgumentParser(
        prog='LinearRegression',
        description='Randomly generates data and fits a linear regression model using Spark MLlib.'
    )
    parser.add_argument('--s3bucket', required=True)
    args = parser.parse_args()
    s3_path = args.s3bucket
    # Part 1: Generate Fake Data and Save to S3
    # Create a DataFrame with fake data
    fake_data = spark.createDataFrame([
        (1, 0.5, 1.2),
        (0, 1.0, 3.5),
        (1, 2.0, 0.8),
        # Add more rows as needed
    ], ["label", "feature1", "feature2"])
    # Perform necessary data transformations
    # (e.g., feature engineering, handling missing values, etc.)
    # ...
    # Prepare the data for model training
    # Assuming the features are in columns "feature1" and "feature2"
    assembler = VectorAssembler(inputCols=["feature1", "feature2"], outputCol="features")
    data = assembler.transform(fake_data)
    # Save the processed data to S3
    data.write.parquet("s3a://" + s3_path + "/path_to_processed_data.parquet")
    # Stop the Spark session
    spark.stop()

training_script.by

import argparse
# Import necessary libraries
from pyspark.sql import SparkSession
from pyspark.ml.feature import VectorAssembler
from pyspark.ml.classification import RandomForestClassifier
from pyspark.ml.evaluation import MulticlassClassificationEvaluator
if __name__ == "__main__":
    # Initialize Spark session
    spark = SparkSession.builder.appName("TrainingApplication").getOrCreate()
    parser = argparse.ArgumentParser(
        prog='LinearRegression',
        description='Randomly generates data and fits a linear regression model using Spark MLlib.'
    )
    parser.add_argument('--s3bucket', required=True)
    args = parser.parse_args()
    s3_path = args.s3bucket
    # Part 2: Read Data from S3, Model Training, and Save Model to S3
    # Read the processed data from S3
    processed_data = spark.read.parquet("s3a://" + s3_path + "/path_to_processed_data.parquet")
    # Split the data into training and test sets
    train_data, test_data = processed_data.randomSplit([0.7, 0.3], seed=42)
    # Initialize the classification model (e.g., RandomForestClassifier)
    classifier = RandomForestClassifier(labelCol="label", featuresCol="features")
    # Train the model on the training data
    model = classifier.fit(train_data)
    # Save the trained model to S3
    model.write().overwrite().save("s3a://" + s3_path + "/path_to_saved_model")
# Stop the Spark session spark.stop()

evaluation_script.py

import argparse
# Import necessary libraries
from pyspark.sql import SparkSession
from pyspark.ml.feature import VectorAssembler
from pyspark.ml.classification import RandomForestClassifier, RandomForestClassificationModel
from pyspark.ml.evaluation import MulticlassClassificationEvaluator
if __name__ == "__main__":
    # Initialize Spark session
    spark = SparkSession.builder.appName("EvaluationApplication").getOrCreate()
    parser = argparse.ArgumentParser(
        prog='LinearRegression',
        description='Randomly generates data and fits a linear regression model using Spark MLlib.'
    )
    parser.add_argument('--s3bucket', required=True)
    args = parser.parse_args()
    s3_path = args.s3bucket
    # Part 3: Load Model from S3 and Perform Evaluation
    # Load the saved model from S3
    loaded_model = RandomForestClassificationModel.load("s3a://" + s3_path + "/path_to_saved_model")
    test_data = spark.read.parquet("s3a://" + s3_path + "/test_data.parquet")
    # Make predictions on the test data
    predictions = loaded_model.transform(test_data)
    # Evaluate the model's performance
    evaluator = MulticlassClassificationEvaluator(labelCol="label", predictionCol="prediction", metricName="accuracy")
    accuracy = evaluator.evaluate(predictions)
    # Print the evaluation result
    print("Accuracy:", accuracy)
    # Stop the Spark session
    spark.stop()
  • An SNS topic where we’ll send our errors, if they happen
  • A ServiceRole for EMR
    – I used the EMR_DefaultRole, that AWS provides for us 
  • An IAM Role for our EMR clusters (I called it blog-emr-job-role)
    – This will need the following trust policy:
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "Service": "ec2.amazonaws.com"
            },
            "Action": "sts:AssumeRole"
        }
    ]
}

                 – And the following permissions:

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Resource": "*",
            "Action": [
                "cloudwatch:*",
                "ec2:Describe*",
                "s3:*"
            ]
        }
    ]
}
  • And an IAM Role for our Step Function
    – With the following trust policy:
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "Service": "states.amazonaws.com"
            },
            "Action": "sts:AssumeRole"
        }
    ]
}

                 – And the following permissions:

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Action": [
                "states:Create*",
                "states:Describe*",
                "states:StartExecution",
                "elasticmapreduce:RunJobFlow",
                "elasticmapreduce:*",
                "states:List*",
                "iam:PassRole",
                “sns:*”
            ],
            "Resource": "*",
            "Effect": "Allow"
        }
    ]
}

Once we have all of these elements ready, we’ll need to replace them in our StepFunction definition, which I share below.

At a high level, we define three steps for each stage of our workflow: cluster creation, code execution and cluster termination. We also define a “catch-all” step that will alert via SNS if there is an error on any of the code execution steps. Something interesting to highlight is that StepFunctions have a predefined step type (createCluster, addStep and terminateCluster) for each of the steps we need, which reduces significantly the amount of code we need to write for our orchestration.:

{
  "Comment": "EMR Data Processing Workflow",
  "StartAt": "DataProcessing",
  "States": {
    "DataProcessing": {
      "Type": "Task",
      "Resource": "arn:aws:states:::elasticmapreduce:createCluster.sync",
      "Parameters": {
        "Name": "DataProcessingCluster",
        "ReleaseLabel": "emr-6.4.0",
        "LogUri": "s3://<YOUR_S3_BUCKET>/logs/",
        "Instances": {
          "InstanceFleets": [
            {
              "Name": "Master",
              "InstanceFleetType": "MASTER",
              "TargetOnDemandCapacity": 1,
              "InstanceTypeConfigs": [
                {
                  "InstanceType": "m5.xlarge"
                }
              ]
            },
            {
              "Name": "Core",
              "InstanceFleetType": "CORE",
              "TargetOnDemandCapacity": 1,
              "InstanceTypeConfigs": [
                {
                  "InstanceType": "m5.xlarge"
                }
              ]
            }
          ],
          "KeepJobFlowAliveWhenNoSteps": true,
          "TerminationProtected": false
        },
        "Applications": [
          {
            "Name": "Spark"
          }
        ],
        "ServiceRole": "arn:aws:iam::<YOUR_ACCOUNT_ID>:role/EMR_DefaultRole",
        "JobFlowRole": "arn:aws:iam::<YOUR_ACCOUNT_ID>:instance-profile/<YOUR_EMR_ROLE_(blog-emr-job-role)>",
        "VisibleToAllUsers": true
      },
      "ResultPath": "$.DataProcessingClusterResult",
      "Next": "DataProcessingStep"
    },
    "DataProcessingStep": {
      "Type": "Task",
      "Resource": "arn:aws:states:::elasticmapreduce:addStep.sync",
      "Parameters": {
        "ClusterId.$": "$.DataProcessingClusterResult.ClusterId",
        "Step": {
          "Name": "DataProcessingStep",
          "ActionOnFailure": "TERMINATE_CLUSTER",
          "HadoopJarStep": {
            "Jar": "command-runner.jar",
            "Args": [
              "spark-submit",
              "--deploy-mode",
              "cluster",
              "s3://<YOUR_SR_BUCKET>/preprocessing_script.py",
              "--s3bucket",
              "<YOUR_SR_BUCKET>"
            ]
          }
        }
      },
      "ResultPath": "$.DataProcessingStepResult",
      "Next": "Terminate_Pre_Processing_Cluster",
      "Catch": [
        {
          "ErrorEquals": [
            "States.ALL"
          ],
          "Next": "HandleErrors"
        }
      ]
    },
    "Terminate_Pre_Processing_Cluster": {
      "Type": "Task",
      "Resource": "arn:aws:states:::elasticmapreduce:terminateCluster.sync",
      "Parameters": {
        "ClusterId.$": "$.DataProcessingClusterResult.ClusterId"
      },
      "Next": "Training"
    },
    "Training": {
      "Type": "Task",
      "Resource": "arn:aws:states:::elasticmapreduce:createCluster.sync",
      "Parameters": {
        "Name": "TrainingCluster",
        "ReleaseLabel": "emr-6.4.0",
        "LogUri": "s3://<YOUR_SR_BUCKET>/logs/",
        "Instances": {
          "InstanceFleets": [
            {
              "Name": "Master",
              "InstanceFleetType": "MASTER",
              "TargetOnDemandCapacity": 1,
              "InstanceTypeConfigs": [
                {
                  "InstanceType": "m5.xlarge"
                }
              ]
            },
            {
              "Name": "Core",
              "InstanceFleetType": "CORE",
              "TargetOnDemandCapacity": 1,
              "InstanceTypeConfigs": [
                {
                  "InstanceType": "m5.xlarge"
                }
              ]
            }
          ],
          "KeepJobFlowAliveWhenNoSteps": true,
          "TerminationProtected": false
        },
        "Applications": [
          {
            "Name": "Spark"
          }
        ],
        "ServiceRole": "arn:aws:iam::<YOUR_ACCOUNT_ID>:role/EMR_DefaultRole",
        "JobFlowRole": "arn:aws:iam::<YOUR_ACCOUNT_ID>:instance-profile/<YOUR_EMR_ROLE-(blog-emr-job-role)>",
        "VisibleToAllUsers": true
      },
      "ResultPath": "$.TrainingClusterResult",
      "Next": "TrainingStep"
    },
    "TrainingStep": {
      "Type": "Task",
      "Resource": "arn:aws:states:::elasticmapreduce:addStep.sync",
      "Parameters": {
        "ClusterId.$": "$.TrainingClusterResult.ClusterId",
        "Step": {
          "Name": "TrainingStep",
          "ActionOnFailure": "TERMINATE_CLUSTER",
          "HadoopJarStep": {
            "Jar": "command-runner.jar",
            "Args": [
              "spark-submit",
              "--deploy-mode",
              "cluster",
              "s3://<YOUR_SR_BUCKET>/training_script.py",
              "--s3bucket",
              "<YOUR_SR_BUCKET>"
            ]
          }
        }
      },
      "ResultPath": "$.TrainingStepResult",
      "Next": "Terminate_Training_Cluster",
      "Catch": [
        {
          "ErrorEquals": [
            "States.ALL"
          ],
          "Next": "HandleErrors"
        }
      ]
    },
    "Terminate_Training_Cluster": {
      "Type": "Task",
      "Resource": "arn:aws:states:::elasticmapreduce:terminateCluster.sync",
      "Parameters": {
        "ClusterId.$": "$.TrainingClusterResult.ClusterId"
      },
      "Next": "Evaluation"
    },
    "Evaluation": {
      "Type": "Task",
      "Resource": "arn:aws:states:::elasticmapreduce:createCluster.sync",
      "Parameters": {
        "Name": "EvaluationCluster",
        "ReleaseLabel": "emr-6.4.0",
        "LogUri": "s3://<YOUR_SR_BUCKET>/logs/",
        "Instances": {
          "InstanceFleets": [
            {
              "Name": "Master",
              "InstanceFleetType": "MASTER",
              "TargetOnDemandCapacity": 1,
              "InstanceTypeConfigs": [
                {
                  "InstanceType": "m5.xlarge"
                }
              ]
            },
            {
              "Name": "Core",
              "InstanceFleetType": "CORE",
              "TargetOnDemandCapacity": 1,
              "InstanceTypeConfigs": [
                {
                  "InstanceType": "m5.xlarge"
                }
              ]
            }
          ],
          "KeepJobFlowAliveWhenNoSteps": true,
          "TerminationProtected": false
        },
        "Applications": [
          {
            "Name": "Spark"
          }
        ],
        "ServiceRole": "arn:aws:iam::<YOUR_ACCOUNT_ID>:role/EMR_DefaultRole",
        "JobFlowRole": "arn:aws:iam::<YOUR_ACCOUNT_ID>:instance-profile/<YOUR_EMR_ROLE_(blog-emr-job-role)>",
        "VisibleToAllUsers": true
      },
      "ResultPath": "$.EvaluationClusterResult",
      "Next": "EvaluationStep"
    },
    "EvaluationStep": {
      "Type": "Task",
      "Resource": "arn:aws:states:::elasticmapreduce:addStep.sync",
      "Parameters": {
        "ClusterId.$": "$.EvaluationClusterResult.ClusterId",
        "Step": {
          "Name": "EvaluationStep",
          "ActionOnFailure": "TERMINATE_CLUSTER",
          "HadoopJarStep": {
            "Jar": "command-runner.jar",
            "Args": [
              "spark-submit",
              "--deploy-mode",
              "cluster",
              "s3://<YOUR_SR_BUCKET>/evaluation_script.py",
              "--s3bucket",
              "<YOUR_SR_BUCKET>"
            ]
          }
        }
      },
      "Next": "Terminate_Evaluation_Cluster",
      "Catch": [
        {
          "ErrorEquals": [
            "States.ALL"
          ],
          "Next": "HandleErrors"
        }
      ]
    },
    "Terminate_Evaluation_Cluster": {
      "Type": "Task",
      "Resource": "arn:aws:states:::elasticmapreduce:terminateCluster.sync",
      "Parameters": {
        "ClusterId.$": "$.EvaluationClusterResult.ClusterId"
      },
      "End": true
    },
    "HandleErrors": {
      "Type": "Task",
      "Resource": "arn:aws:states:::sns:publish",
      "Parameters": {
        "TopicArn": "arn:aws:sns:us-east-1:<YOUR_ACCOUNT_ID>:YOUR_SNS_TOPIC",
        "Message": "An error occurred - The last EMR cluster was not terminated to allow further analysis of the error"
      },
      "End": true
    }
  }
}

Some interesting aspects we can see – for example:

"$.EvaluationClusterResult.ClusterId"

makes reference to the step above:

"ResultPath": "$.EvaluationClusterResult", 

which allows us to create the cluster in one step, obtain its id and then use it in another step.

If we wanted to do more complex things, the idea is the same – we have to reference it in the result path or output path. More details about this here

Going back to our stack creation – once we’ve replaced the values, we can author a new Step Function. You can use this link or navigate to the StepFunctions console, select “Create State Machine” and select the option “Write your workflow in code”.

Both options will open the screen we see below, that includes a basic “Hello World” application. You will replace the code in the definition, and refresh the graph.

Once you’ve refreshed the graph, you should see the following:

The next steps are to name your Step Function, and assign it a role (the one we had created above would be ideal) – and create it with the button “Create State Machine”. Allow some minutes for it to be completed, and you’re ready to go! 

You’ll find a button to “Start Execution” – which will trigger your steps. Once it starts executing, the flowchart will update with the different statuses – blue for steps in progress, green for successful steps, orange for caught errors, and red for unexpected errors. 

It’s important to highlight that the EMR clusters will be available via the usual EMR console – which simplifies the process monitoring for people who are familiar with it. The fact that it was triggered via StepFunctions is completely transparent to the end users. 

Going back to our Step Function and error handling, in the image below we can see two different errors. The step “TrainingStep” caught an error (which is part of the expected workflow), but the step “HandleErrors” failed. This allows us to easily find what failed and where.

On this other image, we can see that the workflow failed, but the error handling succeeded. 

This is important to highlight because caught errors are part of our workflow and we have prepared for them (e.g. with alerting). Unexpected errors might leave our work in an incomplete state that requires manual intervention.

Conclusions and next steps

Well, thanks for making it this far and I hope you’re excited to try this out!

We can see that spinning up a set of EMR clusters is quite easy with Step Functions, and it allows us to leverage the power or EMR in a simple and straightforward way. As a final recommendation, I’d like to highlight some of the best practices for using these two services together:

  • Separate cluster creation and job execution
  • Parametrize your steps
  • Handle failure scenarios gracefully
  • Include log destinations
  • Optimize cluster sizing
  • Be careful with security and access control

With all of this in mind, onwards to your own experiments! And don’t hesitate to write with any issues or questions – happy to help! 

Happy coding! 🙂

Stay ahead of the curve on the latest trends and insights in big data, machine learning and artificial intelligence. Don’t miss out and subscribe to our newsletter!

]]>
Trends to keep an eye out for in 2023 after AWS re:Invent. https://www.montevideolabs.com/2023/02/16/trends-to-keep-an-eye-out-for-in-2023-after-aws-reinvent/ Thu, 16 Feb 2023 16:52:10 +0000 https://www.montevideolabs.com/?p=2734

If you follow our social media you might have realized that our team participated in the latest edition of AWS re:Invent and what an incredible way to finish the year it was! It’s an amazing experience – and an exhausting one! Five days packed with stimulating talks, motivating keynotes, thought provoking peers from different industries and meeting thousands of people from all around the globe.

Now that 2023 has started, and our calendars are back to normal, I wanted to take the time to reflect upon some of the trends I’ll be keeping an eye out for this year, especially after re:Invent.

Data, Big Data

First things first, and something we already knew at Montevideo Labs: Data is the future, but also the present! It was clearly a central part of Adam Selipsky’s Keynote, where he made some impressive announcements. Data is everywhere, and there are a wide variety of sources, which keep growing! AWS is looking to bring a solution through Data Zone – a new data management service that allows us to view and manage data across different sources. Furthermore, Data Zone will allow us to centrally govern data access – a frequent headache for modern organizations with a myriad of data sources and data consumers that only keep growing. The diagram below shows the different components it includes, and how they integrate with different tools. So, we’ll be looking forward to start integrating Data Zone on our upcoming projects and leveraging its power!

                                                                                 Source: https://aws.amazon.com/datazone/

Everywhere AI

Something we are all experiencing – AI is becoming more and more ubiquitous. Entire industries are being disrupted by AI – even sports and health. I attended a session where a speaker used IoT technology to monitor his blood sugar, and we have clients that are taking the world of sports by storm. We will continue to see how AI affects our day-to-day life. As developers, it’s also interesting how the costs for using ML continue to drop. AWS released EMR serverless, which enables us developers to focus on our algorithms and simplifies the managing of clusters, while also keeping costs low. Another aspect that helps us are the large number of “pre built” and “low-code” solutions presented, aimed at reducing toil and letting us developers focus on adding business value. SageMaker Data Wrangler and AWS Comprehend are some of the ones we have played with, and we believe we’re just scratching the surface. 

No time to wait – async is the way!

In the blue pill/red pill universe that Dr. Werner Vogels (AWS Vice President and Chief Technology Officer) presented, we reflected about how the world is asynchronous – and that this is awesome! Montevideo Labs has been using microservices architectures for years now, and all the different event patterns are our bread-and-butter, but it’s amazing how new features promise to impact our lives. Some of the new features for Lambda such as LambdaSnapStart were announced on the more hardware-oriented Monday session, but new features such as EventBridge Pipes promise to simplify producer/consumer integrations and speed up development. 

All the help we can get

AWS has been focused in helping developers make better use of their time, and better use of their tools. The newly announced Amazon CodeCatalyst and Application Composer are only in preview, but the idea of having blueprints or to be able to “drag-and-drop” services is amazing – allowing developers to focus on coding, and DevOps to foster best practices. Along these lines, new “Serverless” have emerged, such as Amazon OpenSearch (formerly ElasticSearch) Serverless and EMR Serverless. These speed up development work, while also reducing DevOps work. 

Start me up!

A group of players in the AWS ecosystem that were particularly pampered at the conference were Startups. There are dedicated teams, support and account managers that specialize in this great group of companies. As Startups ourselves, we had the pleasure of being invited to a session with Ben Horowitz where he reminded us of our purpose: “What is the thing that we are going to do that is going to make a difference?” Having a large group of startup founders, he decided to focus not on the economy, not on business plans, but on reminding us all of the importance of culture. This tells us we’re on the right track. As 2023 stretches its wings, it’s important to remember that AWS has great opportunities for Startups, such as funding, special programs like JumpStart and great partners like us! 

Howdy partner!

Partnerships are becoming more important every day. As partners, we have access to resources and opportunities that make us all better professionals. At the conference, we met with some members of the AWS Partner team who showed us different ways to help our customers with Expertise, Technical Support, Credits, Proofs of Concept or Reviews. Reach out if you’d like to learn more!

Interested in exploring the latest releases from AWS?  As AWS Partners our team at Montevideo Labs has extensive experience with AWS services at scale. Contact our team at info@montevideolabs.com to learn how we can help you on your cloud journey.

 

 

Stay ahead of the curve on the latest trends and insights in big data, machine learning and artificial intelligence. Don’t miss out and subscribe to our newsletter!

]]>
My main takeaways from Peopleware: Productive Projects and Teams – Timothy Lister and Tom DeMarco https://www.montevideolabs.com/2022/06/29/my-main-takeaways-from-peopleware-productive-projects-and-teams-by-timothy-lister-and-tom-demarco/ Wed, 29 Jun 2022 18:02:25 +0000 https://www.montevideolabs.com/?p=1754 I’ve always found New Year’s resolutions to be a very strange thing. They weren’t something that was frequent in my family, and I have never met anyone who made one, and actually stuck to it. However, I decided to make one myself this year: read 52 books in 2022. 52 – you’ll have realized if you’re a numbers’ fan as I am – is not a random number: it’s roughly one book per week (it’s pretty fun to tell people the number and see if they realize the reason or not). There’s no particular genre, nor length. I’ve read skinny books that let me rest a bit, and thick books that I need to plow through. Novels and memoirs, business and self help books, everyone’s welcome.

My second goal – slightly more altruistic – is to try and share some learnings from the ones I particularly like. So, this first post is going to be related to a book that our CEO Maximo recommended (and lent): Peopleware: Productive Projects and Teams, by Timothy Lister and Tom DeMarco. Fun fact, there were two copies in the office, so my husband and I were reading it at the same time (we started it at the same time… but it wasn’t a race. Else I would’ve won. Duh).

With such a title, I have to admit I expected a pretty boring book. I looked at it sitting on my nightstand, looking so smug, half bitten by Maximo’s dog… it didn’t look too tempting. What can I say, I’m a living cliche: judging a book by its -bitten- cover.  240 pages later, I’m reminded of the truth of that saying. I half-filled it with markers and post-its to discuss it with the team, and I actually laughed out loud at some of the truths I found inside its pages. Here are some of my  takeaways:

Project Management

The first concept that made me laugh was Parkinson’s Law: “work expands to fill the time allocated for it”. Written by a humorist, I thought about how many times we’ve said: “there’s no time for this”, or, “since we have the time, let’s do this or that”. As a developer, lead, manager, sometimes it feels like our teams will always need more time for a certain project, so is it possible that the allotted time influences how we do things?

Engineering Management

As I’ve transitioned from more technical to managerial roles, I’ve often wondered what’s the real job description of a manager. I loved the definition given in the book: “The manager’s function is not to make people work, but to make it possible for people to work”. Removing obstacles, helping teams jell (another hugely important concept – how to help teams work well together, cooperate, and become a unit), identifying key players and their role, allowing people to attain a state of flow (that magic moment where you’re in a deep, almost meditative involvement, and are able to make code “flow” from your fingers – you know what I mean).

Another aspect of management is to help people improve. This can be providing opportunities for private self-assessment, to conducting thorough 1x1s with teams. Maybe even having in-house aptitude tests for people to self-assess objectively, and compare with what they thought of themselves beforehand. This ties hand-in-hand with another concept that the book mentions: everyone wants to work in teams that do things with high quality. Helping people improve also works towards that goal.

Finally, I loved this phrase: “the ultimate management sin is wasting people’s time”. 

Team diversity

Another moment that made me laugh was when the authors said: “The ‘class picture’ you take of your next project team is likely to show something that looks more like a United Nations task force than the kind of single-culture group that our fathers and grandfathers managed”. That made me reflect about the importance of diversity, and how that helps teams. The authors also make a point for diverse teams, with different genders, orientations, backgrounds… It’s something we strive for, and it’s great to read that our instincts are correct.

Hiring, training and coaching

Our interview process often includes an interview with future team members. On the one hand, this goes against favoring “flow”, because it adds one more meeting to our folks. On the other hand, it’s mentioned as one of the best ideas for hiring: testing how (potential) future colleagues will work together. It’s a beautiful thing to witness great work relationships being born in an interview, with interviewer and interviewee jamming together. 

Once they’re hired, as part of my role as VP of Engineering, my job is to train the new hires (and sometimes retraining people with varied backgrounds). When I mentioned that to my friends, they were shocked – almost as if they thought that it was a waste of my valuable VP time. However, as part of a section called “The Right People”, the authors mention the importance of training and retraining people, and how at Hitachi Software, the chief scientist has as his principal function the training of new hires – and they’re not the only ones. So, if we’re crazy, at least we’re in good company. 

After being trained and successfully deployed into teams, it’s important to continue coaching our team. It can be both peer-coaching or it can be someone in higher management, but it’s important to have someone that holds us accountable to our long term personal and professional growth. Again, similar to 1x1s, it takes time. But things that are worthwhile usually do. Having teams that build satisfying communities, and care about coaching, career progress and learning are the ones that build satisfying communities and tend to keep their people.

Meetings

It’s important to try to have good meeting hygiene: only the necessary people, try to make them as short as possible, and be able to know when they’re done (decision made, agreement reached, etc). Having too many people (who often don’t speak) just to FYI them, makes no sense. If you have a status meeting, try to have it adding value to all attendees, not just the manager. And be very careful with recurring meetings: often they are not necessary (except ceremonies).Finally, if we have a “working meeting” – what does it say of the rest? 

 Change and learning

Learning is a critical improvement mechanism – those who don’t learn can’t expect to prosper. Often, middle managers are the ones proposing changes and the ones that have the chance to be close enough to the action to learn from mistakes. It’s important to be very attentive to what they say, and to empower them to drive change. Also, change will often have hard detractors and lukewarm supporters – but also people who are believers but questioners: managers should try to identify them, and try to bring them on board. They’re the best chance for the project to succeed.

Chaos

Usually, engineers tend to have problem solving minds, and that makes us value a little bit of chaos: it’s what gives us adrenaline. Working on problems that are already solved, and we can only have limited impact… kinda crushes our soul. So it’s a good thing to try and have some space for chaos (crazy projects, or proofs of concept where we can run wild), because they make us feel alive and challenged.

All in all, this was definitely not a poolside book for January. The thinnest book of 2022 so far, and the one that took me the longest to read, because it made me stop, reflect, highlight… I couldn’t read more than two chapters per sitting – which is pretty bad for my yearly goal, but shows that sometimes, less is actually more! 


]]>
Lessons from software development in times of COVID-19 https://www.montevideolabs.com/2020/04/09/lessons-from-software-development-in-times-of-covid-19/ Thu, 09 Apr 2020 15:05:14 +0000 http://www.montevideolabs.com/ml-ns/?p=205 Working from home is seen like a great idea these days, since staying at home seems to be the most effective way to combat the spread of the Covid-19 pandemic. As someone who works in the software industry, my work isn’t too disrupted by why or where I work, as long as there’s a decent internet connection. But why is that? A simplistic answer would be that I don’t really need any other materials, or that we’re used to being in front of a screen 24/7 (and that we’re already pretty socially awkward!). However, I believe that the best advantage that software has for working from home is the way we organize our work, using Agile methodologies. The point of this article is to share some ideas, tips and tools, for you to try. 

Disclaimer: of course not everyone can apply these. Or at least not directly. But some ideas can be applied anyway. For instance, having a communication channel with your coworkers, other than WhatsApp applies to most industries.

Organizing the work

So, what is Agile? Agile is a way to organize projects. It is characterized by the division of tasks into short iterations of work and frequent reassessment and adaptation of plans. If we need to do a task, the agile way is to break it down into smaller, simpler parts, and build the ones that add value first. As we do this, we can always find a way to make things better and more productive. 

From these principles (that I’ve disrespectfully summarized in three lines), several methodologies and sets of practices have been born. In this case, I’m going to make reference to Scrum, since it’s the one I like the most, and the one I’ve worked with the most. Scrum emphasizes on daily communication and the flexible reassessment of plans that are carried out in short, iterative, phases of work. Some of the ceremonies included in Scrum are the daily meeting, sprint planning and backlog grooming.

What is a daily meeting, and why should my team do it? The daily meeting is a quick, five-minute video-call meeting in which every team member shares what they did the previous day, what they plan on doing that day, and if they’re blocked. It has several advantages. First, it requires the entire team to get up, get dressed and show up, especially if you enforce the use of webcams. It also helps reduce repeated work, because it’s simple to realize if two people are working on the same task. When working from home, there’s no more small talk over the watercooler about what you’re doing, so keeping everyone in the loop is more complicated. It promotes team cooperation and reduces blockers, since whoever is blocked can ask for help. It’s also a great opportunity, if you have kids, to let them join in and say “Hi”. We can all use an extra dose of cuteness, and it can help them understand that you are working. Just like bring-your-kid-to-work-day, but with less breakables and a mute button!

In order for everyone to have tasks, it’s necessary to organize the pending work. The Scrum way of doing this is keeping a board. I personally recommend Trello, since it’s free of charge and easy to use. (Link here: https://trello.com/guide/create-a-board). With a Trello board, everyone on the team can add in tasks to the “Backlog” list. It can be an actual task, or an idea to analyze with the rest of the team. Each task has a card that people can write comments on, add links, pictures or any other piece of useful information for that task. For example, someone needs to circle back with a customer? Add a ticket with the description, and someone will grab it in the next planning session.

Once we have tasks in the backlog, how do we organize the tasks? First, we should have a weekly 30-minute session to organize the pending tasks. We should keep at the top the tasks that are more important according to our priorities. We should also consider which tasks block other tasks, and which ones add value. For instance, creating a wireframe can be considered less important, but it adds value, since it lets us show an intermediate product to the customer, and validate our design. Thanks to this quick meeting, we can have an organized roadmap, and we don’t lose good ideas, since they are all available on the board.

Finally, we should run weekly or bi-weekly planning as a team. With an organized backlog, knowing which tasks are more pressing is simpler. Each team member should move the tasks that they consider that can be finished in the allocated time. Those tasks are assigned to each team member, and a small picture/name is attached to the ticket, so we can see at a glance what everyone is working on. As the tasks are done, they navigate from “This week” to “Ready for review”, and finally they get moved to the “completed” column. This way, any team member can see what the rest are up to at a glance. 

Team communication

The other important component of a successful distributed workforce are good communication tools. Personally, I believe that companies should use specific communication tools, like Slack or Zoom, instead of phone calls and iMessage/WhatsApp, because it helps create a barrier. If everyone communicates via WhatsApp, we’re all available 24/7, and that is the express way to burn-outville. Also, while working, you only need to have your work related tools open, thus reducing the chances of getting distracted by the other million or so chats going on.

What tools should we use? First, a good messaging tool. I personally recommend Slack. Why? First, it’s free, and easy to use and customize. It lets you have one-on-one chats, multiple-person conversations, private rooms with several people and open rooms where everyone can join in and see what’s being talked about. This helps organize the communication channels immensely. 

A personal recommendation is to use all the versatility that Slack brings. For instance, create a channel with jokes, or other fun stuff. Why? Because it helps people decompress, and share a laugh, like they would if they were sharing an office. The daily jokes and stories we tell every day are lost in the home office environment, so organizations should try to keep the office spirit alive. Also, when chatting, it’s pretty easy to misinterpret someone else’s tone. Sharing a joke is an easy way to remind everyone that all is ok. Finally, having a single place to share the fun stuff keeps the work channels clear of noise.

Another fun idea is to look for shared interests, like games, or dogs, or cooking. Create a channel for this, and let people join in and have fun. This creates the bonds that home-officing separates. Lots of offices have a pets channel, where they share their fur babies’ best pictures, and people from different places and teams get together and share a laugh. 

Which other tools are important? I recommend choosing one video conferencing tool in the organization and learning all its tricks. I like Zoom and Google Hangouts Meet, because they allow screen sharing, have intuitive, easy-to-use interfaces and don’t require your users to add contacts. You simply send your meeting link and people can join in. As another alternative, Whereby is another good tool for simplicity. Remember to keep a neutral background behind you, and let the rest of the household know you’re in a videoconference. It can save you a lot of awkward moments 🙂

So, all in all, keep in mind that there is no better tool for cooperation than good humor and a sense of teamwork; especially to support each other through tough times. And remember, this too shall pass!

]]>
The Quarantine Diaries https://www.montevideolabs.com/2020/03/09/quarantine-diaries/ Mon, 09 Mar 2020 20:09:40 +0000 http://www.montevideolabs.com/ml-ns/?p=1 These recent times have been atypical for everyone. Empty streets, closed stores and stay-at-home campaigns have changed our routines, with everything and everyone compelling us to stay home and prevent the spread of COVID-19 .

When speaking with a friend who works as a teacher about this new “home office” routine, and how we’re both trying to adapt to it, she inquired: “But… How is this any different to your normal days? You only need your computer to work, and online meetings are part of everyday routine, right? Isn’t this the same?”

That simple question prompted me to analyze my experience and to write it down. Because it’s not the same. And it’s not just because my office has the best coffee maker in the world, and now I’m settling for instant coffee (send help!). It’s not even because I miss my office friends and lunchtime jokes. I simply miss that feeling of “being at work”, of hearing someone type aggressively when there’s a frustrating bug, or that cheer when that rebellious test finally passes. I miss laughing when one of my friends shouts “shoot, I had a meeting five minutes ago!”, or having someone to duck test my code. I’ve built a work life that I enjoy. I simply love my day to day, and I feel homesick… or should I say officesick?

Since a lot has been said and written about best practices when working from home, I don’t think I can add anything new. However I’d like to add the perspective of having two people working from home in a small apartment. Santiago and I are both software engineers who work at the same company, for two different customers, so I’d like to share some of our tips (and mistakes!) from the time spent so far in confinement. 

Create a comfortable and distraction free place, and try to use it only while working

Often, we feel tempted to say: I’m working from home, I’ll work from the bed, or sofa. Trust me on this one: not the greatest idea. Two hours after I started typing from the couch, my lower back reminded me that I’m no longer twenty (and I’ve never been flexible).

So now,  I’m working at the dining room table (that has been reconverted into my personal desk). I use an extra screen, keyboard and trackpad, so I take up half of the table. I love working with natural light, so this setting is ideal for me, since I have a nice view and lots of light. One disadvantage is that our dining room chairs aren’t ergonomic, but I’ve solved it with a pillow and occasional stretches.  

Santiago prefers his desk chair, so he set up camp in the flat’s office, where he has created an “office vibe”. He doesn’t need as many gadgets as I do, and needs to be able to close his door, since his normal day includes a million meetings (or so it seems from the outside).

So, find a spot where you can be the least distracted and that you can “take over” (for example, not in front of the TV!).

Regarding the computer and screen arrangements, try to position your screens in a way that doesn’t hurt your back. Some recommendations are:

  • Position your monitor so that the top of the screen is at eye level.
  • Your eyes should be looking slightly downward when you’re viewing the middle of the screen.
  • The screen should be about an arm’s length distance.
  • Adjust the tilt to reduce the glare

Divide up the tasks evenly

We usually try to share the chores, but during the day it can be difficult with meetings and calls. So, if one has some extra time before lunch, that person can prepare the meal or set the table, but then the other should clean (if meetings and schedules allow). It’s good for everyone, because both get the chance of having an “active pause” and a break from sitting in front of the computer. Also, as an extra bonus, nobody carries the full weight of chores.

Try to cook before/after work

A fresh cooked meal is always tempting. However, cooking takes time, and even if you have flexible hours, it’s very likely that your teammates are counting on you to be online and available during office hours. What we’re trying to do is to plan the weekly meals ahead, go to the supermarket once a week – early in the morning to avoid long queues – and cook in the evenings. That way, we only need to heat up lunch, and we can enjoy some sofa time before going back to work. 

It’s also a good time to try and hone some skills.  For instance, when cooking, I usually avoid cutting meat because I’m obsessed with cleaning it perfectly, and I take forever. Now, I can try and do it, while he chops the vegetables. His chopping skills have improved greatly, and we have lots of fun (and I no longer cry over onions!). Another plus, I’m getting to overcome my disgust when handling raw chicken, which is usually his task. We’re nowhere near MasterChef, but we still haven’t gotten ourselves poisoned. Yet. 

Know what works for you

In spite of what mainstream media tells us, not all software engineers work with their headphones on from 9 to 5. I, for instance, find it uncomfortable to isolate completely from the world around. So, I play some background music using a small speaker, while still hearing noises around me. I’m considering it implementing this for our small pod back at work! 

Another myth to debunk; not everyone is an early bird, not everyone is a night owl, but everyone has their preferences. If you work with people in other timezones, you can try and find a compromise. For instance, you can start working a couple of hours before the rest, and clock out some hours earlier than the rest. 

I prefer to sit at my desk around 9.00 a.m., before my Slack starts getting notifications. I’m an early bird, and having an hour or two by myself allows me to focus. However, since Santiago enjoys afternoon work, we’re usually working until 6:30 or 7.00 p.m., so logging in before 9.00 would mean working longer hours, which isn’t always the best. 

Disconnect your head when you log out

Not having other plans in the day (no gym, no meeting friends, no family get-togethers) can make us forget about the time. Also, since your desk is just steps away, it can be tempting to continue working “after hours”, especially if there’s a task that we’re craving to finish, or that is highly engrossing. However, try to resist. Taking breaks from work actually boosts your creativity, and helps you tackle the issue with a fresh mind the following day.

We usually try to do something active when we close our laptops (some gym, stretching or going for a quick walk). It usually clears our minds, and lets our minds unwind in an easy way. There are lots of exercise routines online that you can do inside without annoying your neighbors. I personally like Nike’s NTC, which has a wide variety of small routines and it has built in timers. 

Look for shared activities, but also value “selfish time” 

It’s great to create a list of movies to watch, or series to binge on together after work, but it’s also really healthy to do something that you enjoy on your own. It can be reading, gaming, exercising, or watching a show that the other person doesn’t love. It’s fun, and it can give you some interesting conversation topics for dinner! We’re not sure how long this quarantine is going to last, so we should try to take advantage of this time to do those things we never find the time for.

For instance, I’m finishing an online course that I always postponed, and at nights, I sometimes watch an episode of “This is us”. Santiago likes joining his friends for some gaming sessions, and he has some books that have been collecting dust since our holidays finished. I’m thinking of going back to my yoga days. It doesn’t matter what you do; but do something on your own for a little while. 

All in all, we should try and see this crisis as an opportunity to have some nice family time, to slow down and appreciate all the things we usually take for granted. There’s an interesting reflection going around the internet, that says that this disease will teach us to be more selfless, more generous and more caring. I certainly hope it does. Meanwhile, stay safe! 

]]>