04 ago
|
Mindtech
|
Buenos Aires
04 ago
Mindtech
Buenos Aires
Descripción de la vacante
DescriptionOur Data Trainer / Data Scientist generates and maintains high-quality datasets across different business domains and ensure that all samples are well-written, technically sounds, and useful for the end users. By researching, and creating targeted content for AI/software developers, QA teams, and field engineers he/she will be an essential component of improving our data-driven solutions and expanding our business offer.Seniority: Ideally we are looking for a senior member who can work independently and bring new creative ideas to the table.Responsibilities- Create data sets that represent customers' data for training of modules, QA team,developers, and field engineers- Develop sensitive data elements extractions for the product and custom ones per customer needsRequirements:
- Extensive experience in developing complex ETL pipelines for data analytics and dataset preparation for machine learning development and QA testing, ideally containing natural language text and patterns- Experience in creating, maintaining, and serving well organised datasets containing both tabular data and unstructured documents- Proficient in Python and industry standard NLP and analytics tools such as pandas, numpy, Gensim, spaCy, NLTK, SQL and NoSQL databases- Meticulous attention for data quality, keen understanding of business needs across different industry domains, and drive for continuous improvement of data-driven solutions- Ability to write high quality modular code in a collaborative environment, and contributing to code reviews within the team- Experience in working with software developers, product managers,
and business stakeholder towards integration of data solutions and refinement of business requirements- Effective communication and ability to write clear and well organized software and data documentationNice to have- Familiarity with text analytics pipelines, and some hands-on experience with training/testing of machine learning models for text classification and entity detection using real-world datasets- Familiarity with web scraping and automated content search and creation through LLMs and prompt engineering (or an interest to learn it)- Familiarity with software development life-cycle, CI/CD pipelines, and MLOps best practicesOther technologies:
- LLMs: previous use on LLMs on a real business case, especially content/data generation is highly appreciated- Cloud Computing platform: Google Cloud and AWSBenefits- Friendly and highly professional atmosphere, laptop or workstation, corporate events.
- Unlimited access to the Well-being platform "Rozumiu" for the employee and their family.
- Paid sick leave days, vacation, and national holiday days.
- Great opportunities for professional growth and advancement.About the projectThe Product has the ability to deliver a very accurate master catalog of sensitive data usage to allow businesses to manage data security/compliance to complement their infrastructure-based security/compliance programs. It is a fully automated solution that covers data in any format, be it structured or unstructured, data-in-motion or data-at-rest, both known or unknown. It covers all aspects of data processing in one place and aggregates that into a master catalog containing all the customers' or employee's information.
📌 Data Trainer - Machine Learning & Nlp (Buenos Aires)
🏢 Mindtech
📍 Buenos Aires