Mario J.
Available now

Mario J.

Senior Data Engineer

19 years of experience

Guatemala 🇬🇹
English (C1)
Interview

Short bio

Over the past 8 years, I have gained ample experience in ETL/ELT, Snowflake, Airflow, and Data Warehouse. I have been responsible for Data Lake Support and creating new ETL/ELT processes within AWS using either Snowflake, Glue, Airflow, dbt, Xplenty, Docker, or Custom Python Scripts with APIs. I have also worked as an Application Architect, which includes over 50 different services across AWS. I needed to create pipelines, organizations, and permissions within AWS Infrastructure to meet business requirements while optimizing costs. My previous work includes over four years as a Data Analyst, creating and supporting BI Tools like Sisense, Quicksight, and Power BI/Tableau. I managed and created dashboards, widgets, and Data Architecture with Sisense and Redshift. I was also the Support Engineer for Sisense Installation and Maintenance. I would describe myself as reliable and able to undertake complex situations with unique solutions. I have over 7 years of experience in the Data sector, over 17+ years of SQL experience, and 20+ years as a Developer. I’m also a mid-QA Automation Engineer, which gives me a better idea of the attention to detail a developer needs to attain to complete development within the time limit and without affecting PROD releases. My other skills include being a Sysadmin in both Windows and Linux environments, a company Web Services AWS, a company 365 Portal, JumpCloud, 1Password, and Slack. I also have DBA knowledge, including creating servers and Data Warehouses. I'm an expert in MS Office and can automate with macros. I’m also a Cybersecurity advocate since I’m always looking for ways to protect data and user permissions. I have been in Agile Development with Jira in Kanban and Scrum for the past 4 years. And I have also created bots to automatize process alerts and configurations with Slack APIs and AWS to help my teammates get prompt feedback from the ETL pipelines.

Tech stack

Main technologies
PythonSQLPandasPySparkJavaVBA.NetAWSAirflowSnowflakeDBTAPIsSingleStoreDruidPostgreSQLMSSQLOracleMySQLRedshiftSisense

Work experience

  • Senior Data Engineer2024 - Present3 y.

    My job has been focused on migrating and building different ETL pipelines to work on, from extracting Information from APIs, as well as adding new reports to existing DAGs for sending data from our Data Warehouse to new customers, I also created the framework for a migration using PySpark since they were having issues moving some heavy-usage processes that required more than 300GB of RAM and they needed to move to a smaller Airflow Server.

    Responsibilities:
    • Responsible for migrating heavy-intensive jobs in their old Airflow installation to AWS Airflow, reducing the server instance type and moving to an EMR solution to leverage the RAM spikes for some jobs (300GB+ RAM).
    • Migrated complex analytics queries from Druid to SingleStore, optimizing query logic to match engine-specific syntax and performance.
    • Refactored expressions into clean, maintainable CTEs, improving performance and readability.
    • Debugged and resolved time zone misalignments between data stored in UTC and reporting in EST/CST/PST, aligning timestamp filters for business logic accuracy.
    • Built an Airflow pipeline to extract paginated API data using AWS Secrets Manager and Pandas, storing results in partitioned S3 paths optimized for Iceberg and Athena. Designed dynamic file naming to support multiple daily ingestions while ensuring query efficiency.
    • Built scalable Airflow DAGs using dynamic task configs, optimized large-scale file processing (8K–22K files) with multithreading, implemented fault-tolerant EMR retries, S3 multipart uploads, and validated SFTP transfers for data integrity.
    • Automated a local tool that processed a file to export to a provider and used different APIs calls to process and monitor its status.
    • Created a Slack bot for notifications of failed DAGs since their current tool (PagerDuty) was lacking information, and with my approach, we would get real-time notifications and exact DAG failures.

    Technologies: AWS, Airflow, Python, Github, SQL, S3, Jira, Confluence, Druid SQL, MySQL, SFTP, APIs, Snowflake, SingleStore, Docker, EMR, PySpark, Slack, PagerDuty

  • Senior Data Engineer2023 - 20241 y.

    My job was building different ETL pipelines, from extracting information from Workday to connecting to financial systems like Adyen. I made different procedures in Snowflake and Python to finish within the deadlines.

    Responsibilities:
    • Responsible for migration of Adyen SFTP process pipeline & migration of Paypal SFTP Process pipeline.
    • Creating and developing new Snowflake Procedures.
    • Creation of lambdas for new data imported from Accounts Receivable Aging Data.
    • Creation of views and procedures for Intacct information and Zuora.
    • I proposed a better way to get Snowflake fails or ORCs (a company Job internal tool) notifications, since most of them did not have a proper message when the process failed.

    Technologies: AWS, Snowflake, Python, Gitlab, Kubernetes, SQL, S3, Jira, Confluence, PostgreSQL, MySQL, Linux, SFTP

  • Senior Application Architect2022 – 20231 y.

    This was a 1-year, time-limited contract to help them set up the foundation of their AWS and several pipelines for their 15 different clients.

    Responsibilities:
    • Creation of 15 different AWS Organizations and ETL pipelines to extract information from current clients with current standards, since most of those integrations were CSV or Excel in Emails or WhatsApp Messages.
    • QuickSight implementation to help boost their BI team.
    • Helped front-end and back-end developers with Git implementation and servers and services for the Data Scientist team.
    • I helped migrate from email to Slack and 1Password implementation as a cybersecurity measure since most passwords were saved in emails or notepads, and a better security standard for their servers and services. I proposed moving to Agile.

    Technologies: AWS, Airflow, Python, Git, SQL, 1Password, Office365, Slack, QuickSight, Jira, Confluence, PostgreSQL, Linux

  • Senior Airflow Support2020 - 20222 y.

    Basically, my job was to make sure all the data from the different pipelines arrived on time for corporate and shareholders meetings since this was the main ETL process. And, to make sure everything was working correctly with the several Airflow DAGs.

    Responsibilities:
    • Responsible for Airflow DAGs support
    • Responsible for fixing DAGs in Non- prod and Prod environments. Creating and maintaining new Airflow operators.
    • Monitoring over 200 DAGs in Non- prod and prod each daily.
    • I suggested new operators and a better way to organize DAGs retries, since some of them were complex and long and restarted from the beginning instead of the previous steps, like the optimization done. I also at request from the Lead Dev optimized the most complex DAG.

    Technologies: AWS, Airflow, Python, Git, Docker, Kubernetes, SQL, Redshift, API, SFTP

  • Senior AWS / Data Engineer2020 - 20222 y.

    On this position I created new ETL processes for extracting data from different APIs like a company, iAuditor, Migration from MSSQL to Redshift, also giving support to other team members with data lake access.

    Responsibilities:
    • Creation of new data pipelines for Airflow
    • Maintaining and expanding current pipelines
    • Permissions access to the data lake to other team members
    • Main git support to other team members
    • AWS support on S3, Docker, Glue, Athena, EMR, Lake Formation, Redshift
    • I helped achieve some optimizations within the airflow operators that the Lead Dev said were not possible, I also suggested moving from Athena to Redshift which boosted the performance on queries.

    Technologies: AWS, Airflow, Python, Git, Docker, Kubernetes, SQL, Redshift, API, SFTP

  • Application Architect / Data Engineer / Data Analyst / SREs2017 – 20225 y.

    CEW is a multi- corporation located in Peru, the project consisted of reading invoices of all its different tenants (approx. 300) and extract the information found in the invoices like invoice number, invoice type, total, items found and around 5 more fields. We received approximately 1.5M files per month, we had both text files and images. The project was 73 tenants behind schedule and the original process took 3 days to set-up a tenant, it also consisted of an archaic way to extract information that it normally needed constant checks and reconfigurations. The project itself was handed to me to revitalize and generate a new pipeline to replace the old process.

    Responsibilities:
    • Everything from Lambdas, EC2 Servers, Redshift, RDS, S3, configurations, Python code was done by me except for the OCR.
    • Final Infrastructure in AWS Consisted of S3, Lambda, RDS, Redshift, EC2, Fleet / Spot
    • Initial infrastructure also had Athena, Glue, S3 Glacier, QuickSight, Data Pipeline as part of the solution.
    • I basically was able to generate the new process within 3 months exceeding the previous system accuracy from around 45% to 95%, The process took 7 seconds from the receipt of files to writing out results to Redshift, I was also able to lower the setup of tenants from 3 days to 15- 20 minutes. It was also more flexible in finding details on the invoices compared to the previous system, which led to fewer reconfigures. We also gained the ability to get the invoice details, that was not impossible to do with the previous system and that was something the client wanted.

    Technologies: AWS [S3, Lambda, RDS, Redshift, Athena, Glue, S3 Glacier, CloudWatch, EC2, Fleet / Spot, Rekognition, Textract, QuickSight, DataPipeline], Python, Regex, Sisense

  • Business Analyst2007-201710 y.

    a company is a startup company with several projects, and I have helped on 25+ projects in Banking, Manufacturing, Food, and General Industries.

    Responsibilities:
    • Optimized packaging planning across Central America, increasing production by 1.4M pounds in a year.
    • Designed cash transportation systems, reducing movement and cutting costs by over 70% across 4 years.
    • Planned and implemented ATM supply systems, reducing cashouts by 95% and increasing uptime from 95% to 99.5%.
    • Built inventory and purchasing solutions, managing 45K+ SKUs across 100+ stores, lowering stock while boosting best-seller sales.
    • Developed ad-scheduling software to streamline and centralize radio spot operations.
    • Developed optimization and planning software across logistics, cash management, and inventory domains.

    Technologies: Excel, VBA, SQL, VB/VBA, VB.NET, R, ASP, Flex, Flash, ActionScript, APIs, VPN

  • Business Analyst, Full Stack Engineer, Data Scientist2007-201710 y.

    a company is a startup company with several projects, and I have helped on 25+ projects in Banking, Manufacturing, Food, and General Industries.

    Responsibilities:
    • Optimized packaging planning across Central America, increasing production by 1.4M pounds in a year.
    • Designed cash transportation systems, reducing movement and cutting costs by over 70% across 4 years.
    • Planned and implemented ATM supply systems, reducing cashouts by 95% and increasing uptime from 95% to 99.5%.
    • Built inventory and purchasing solutions, managing 45K+ SKUs across 100+ stores, lowering stock while boosting best-seller sales.
    • Developed ad-scheduling software to streamline and centralize radio spot operations.
    • Developed optimization and planning software across logistics, cash management, and inventory domains

    Technologies: Excel, VBA, SQL, VB/VBA, VB.NET, R, ASP, Flex, Flash, ActionScript, APIs, VPN

Didn't find the right engineer?All available engineers
Flag icon

U.S.-Based

Discuss Your Project

This is a no-pressure, 30-minute conversation. We will talk through what you are building, identify risks or unknowns, and outline what it would take to do it right.

Certificates

Let's build together.

Talk with a senior engineer about your product idea, architecture, and what it would take to build it.

Upload File