Original job description
THE COMMONS XR would sincerely like to thank the current internship cohort and look forward to the next cohort interns. As such we are seeking unpaid interns who are pursuing Data Engineering professional experiences. Our company is about creating a platform for better social/emotional personal development as well as helping navigate social challenges, through the use of the next generation tools (MetaVerse and AI modeling). If you are selected, you will be asked to answer a number of questions to determine your qualifications and then given a real-life homework assignment to more fully gauge your foundational knowledge (similar to what you would be interning on with projects with us). Please apply quickly as our internship posting generally closes after one (1) Week. You will be working alongside myself (CEO/CTO) and our Data team. Although NOT a requirement, please make sure your Github shows some recent data engineering projects, helps determine what you know.
You must be available Friday AMs (9-11 AM PST) for our standup and simulated immersive experiences. Please do not apply if these times are not available. Preferred start dates are middle of May, preferred end data is middle of Sept, but different times may be allowed. Remember, we can ask about anything on your resume.
Our Environment
SQL, Python, AWS (Postgres, EC2, S3, Sagemaker, Bedrock, Cloudwatch, VPN), BigQuery, Unity, Docker, Kubernetes, GSuite, Microsoft Office Suite, Slack, Git, Bash, C#, JS, SiteGround (Wordpress/REACT)
Responsibilities
The candidate will be handling data from EC2, S3, PostgreSQL, and BigQuery for real-time and batch inference using LLMs and conventional ML though deployments using Bedrock and Sagemaker. The candidate will either create or use custom data pipelines and data flow from EC2 instances. Collaborate with our Unity Engineering and FullStack teams to develop communications architecture between various platforms and services. Learn how to implement data pipelines and communication infrastructure between our service components and improve on them (reliability, supportability, security, etc). Design, implement and document testing and deployment workflows using Github, Bash and cloud-based services. Debug code and automation issues in a complex service architecture. Identify opportunities for improving reliability, scalability and observability of our infrastructure and best practices.
Required Skills
Python
LLM integration knowledge.
Reading and Querying SQL & NSQL data (Postgres preferred)
Cloud Experience (Azure Cloud, AWS, or GCP experience)
Basic knowledge on data pipelines and concepts (ETL & ELT).Fundamental applications for either Large Language Models (LLMs) or Classical Machine Learning Algorithms (DS, MLE)
Data cleaning and curiosity (ALL)
EDA (DS)
Strong communication skills.
Good troubleshooting skills.
Excellent attention to detail.
Preferred Skill
AWS, particularly S3, EC2, Bedrock, Sagemaker in addition to GCP BigQuery.
Strong in data modeling and API creation and documentation.
Technical experience in data infrastructure, cloud computing, storage systems, and distributed computing frameworks in big data scale systems.
Experience with Prompt & AI Engineering, particularly through using an agent SDK
Proficiency in Python Data Science Libraries
(e.g. Pandas, sklearn, Matplotlib, etc.)
Knowledge on what the ideal LLM for each use case is (foundation model, size, hyperparameters, etc.).
Mathematical/statistical foundation with an understanding on how to apply ML concepts.
Visualization experience with Python.
Nice To Have
Data Warehouse tools.
Tool monitoring (especially AWS CloudWatch)
Bash or similar scripting language.
Docker/Kurbenetes experience.
Deep Learning Experience.
*Things you should learn with us: *
Strong communication skills and experience working in a team environment. Structured and detail oriented approach to problem solving. Ability to work with legacy codebase Provide infrastructure setups (how to-clips documentation). Learn new tools and services.