Volver al ranking

CodeCutTech/Data-science

Jupyter Notebookcodecut.ai/blog

Collection of useful data science topics along with articles, videos, and code

data-sciencemachine-learningnatural-language-processingpythondata-visualizationdata-analysisarticlesartificial-intelligencetime-seriesscraping
Crecimiento de estrellas
Estrellas
4.2k
Forks
1.1k
Crecimiento semanal
Issues
4
2k4k
ago 2020jul 2022jul 2024jul 2026
README

📦 We’ve Relocated
The contents of this repository are now hosted at github.com/khuyentran1401/codecut-blog. Please follow the new repo to stay updated.

Data Science Articles from CodeCut

About CodeCut

CodeCut is the platform that helps data scientists stay productive and current by delivering short, practical code examples that highlight modern tools in action.

It's the resource you wish you had when learning a new library—clean, concise, and instantly applicable.

Article Collection

This repository is a curated collection of data science articles from CodeCut, covering topics like MLOps, data management, testing, visualization, and more. Each article comes with practical examples, code repositories, and video tutorials to help you quickly implement these tools and practices in your own projects.

Category Title Article Repository Video
MLOps Goodbye Pip and Poetry. Why UV Might Be All You Need 🔗
MLOps Stop Hard Coding in a Data Science Project – Use Configuration Files Instead 🔗 🔗 🔗
MLOps Poetry: A Better Way to Manage Python Dependencies 🔗 🔗
MLOps Git for Data Scientists: Learn Git through Practical Examples 🔗 🔗
MLOps 4 pre-commit Plugins to Automate Code Reviewing and Formatting in Python 🔗 🔗 🔗
MLOps How to Structure a Data Science Project for Maintainability 🔗 🔗 🔗
MLOps Build Reliable Machine Learning Pipelines with Continuous Integration 🔗 🔗 🔗
MLOps Automate Machine Learning Deployment with GitHub Actions 🔗 🔗 🔗
MLOps How to Build a Fully Automated Data Drift Detection Pipeline 🔗 🔗 🔗
Data Management Tools Version Control for Data and Models Using DVC 🔗 🔗 🔗
Data Management Tools What is dbt (data build tool) and When should you use it? 🔗 🔗 🔗
Data Management Tools Streamline dbt Model Development with Notebook-Style Workspace 🔗 🔗 🔗
Testing Pytest for Data Scientists 🔗 🔗 🔗
Python Helper Tools Write Clean Python Code Using Pipes 🔗 🔗 🔗
Python Helper Tools Introducing FugueSQL — SQL for Pandas, Spark, and Dask DataFrames 🔗 🔗
Python Helper Tools Fugue and DuckDB: Fast SQL Code in Python 🔗 🔗
Python Helper Tools Marimo: A Modern Notebook for Reproducible Data Science 🔗 🔗
Feature Engineering Polars vs. Pandas: A Fast, Multi-Core Alternative for DataFrames 🔗 🔗
Visualization Top 6 Python Libraries for Visualization: Which one to Use? 🔗 🔗
Python Python Clean Code: 6 Best Practices to Make Your Python Functions More Readable 🔗 🔗 🔗
Logging and Debugging Loguru: Simple as Print, Flexible as Logging 🔗 🔗 🔗
LLM Enforce Structured Outputs from LLMs with PydanticAI 🔗 🔗
LLM Run Private AI Workflows with LangChain and Ollama 🔗 🔗
Speed-up Tools Writing Safer PySpark Queries with Parameters 🔗 🔗
Speed-up Tools Narwhals: Unified DataFrame Functions for pandas, Polars, and PySpark 🔗 🔗
Speed-up Tools Eager to Lazy DataFrames with Narwhals 🔗 🔗
Speed-up Tools Scaling Pandas Workflows with PySpark's Pandas API 🔗 🔗

Contributing

If you're passionate about data science and want to share your knowledge about open-source tools for data processing and LLM applications in Python, we'd love to have you contribute!

To contribute:

  1. Create a GitHub issue:
    • Click on the "Issues" tab
    • Click "New issue"
    • Select "Article Topic Suggestion" template
    • Fill in the template with your article proposal
  2. Read our contribution guidelines
Repositorios relacionados
microsoft/ML-For-Beginners

12 weeks, 26 lessons, 52 quizzes, classic Machine Learning for all

Jupyter NotebookMIT Licensemldata-science
88.4k21.6k
apache/superset

Apache Superset is a Data Visualization and Data Exploration Platform

PythonPyPIApache License 2.0supersetapache
superset.apache.org
73.9k17.9k
Asabeneh/30-Days-Of-Python

The 30 Days of Python programming challenge is a step-by-step guide to learn the Python programming language in 30 days. This challenge may take more than 100 days. Follow your own pace. These videos may help too: https://www.youtube.com/channel/UC7PNRuno1rzYPb1xLa4yktw

PythonPyPI30-days-of-pythonpython
68.8k12.8k
scikit-learn/scikit-learn

scikit-learn: machine learning in Python

PythonPyPIBSD 3-Clause "New" or "Revised" Licensemachine-learningpython
scikit-learn.org
66.7k27.2k
keras-team/keras

Deep Learning for humans

PythonPyPIApache License 2.0deep-learningtensorflow
keras.io
64.2k19.7k
pandas-dev/pandas

Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more

PythonPyPIBSD 3-Clause "New" or "Revised" Licensedata-analysispandas
pandas.pydata.org
49.3k20.2k
GokuMohandas/Made-With-ML

Learn how to develop, deploy and iterate on production-grade ML applications.

Jupyter NotebookMIT Licensemachine-learningdeep-learning
madewithml.com
48.8k7.7k
apache/airflow

Apache Airflow - A platform to programmatically author, schedule, and monitor workflows

PythonPyPIApache License 2.0airflowapache
airflow.apache.org
46.2k17.4k
SimplifyJobs/Summer2026-Internships

Summer 2026 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC.

PythonPyPIinterview-preparationinternships
swelist.com
45.5k3.2k
streamlit/streamlit

Streamlit — A faster way to build and share data apps.

PythonPyPIApache License 2.0pythonmachine-learning
streamlit.io
45.3k4.3k
ray-project/ray

Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

PythonPyPIApache License 2.0raydistributed
ray.io
43.3k7.8k
gradio-app/gradio

Build and share delightful machine learning apps, all in Python. 🌟 Star to support our work!

PythonPyPIApache License 2.0machine-learningmodels
gradio.app
43.2k3.6k