MorphSQL
Turn warehouse SQL into notebook-ready pandas or
PySpark — paste or upload a file, convert, then download
.py / .sql to your machine.
- Choose input dialect + output · 2. Paste SQL, load an example, or upload a
.sql/.zip· 3. Convert → preview + download to your machine
Fills SQL, converts, runs sample preview.
| morphsql_pandas .py | 822.0 B ⇣ |
Sample preview
1 | 20 | 1 |
Snowflake → Python (pandas)
92% confidence · output: Python (pandas)
Next steps
- Review the generated Python on the right.
- Check the sample preview (synthetic tables).
- Download the
.pyfile or open the notebook starter cell. - Point
tables['…']at your real DataFrames (read_parquet/read_sql).
What changed
- Source dialect normalized: snowflake → portable SQL → pandas
- WHERE → DataFrame.loc[mask]
- SELECT columns → DataFrame projection
Sample preview · generated pandas on 1 synthetic table(s) → 5 rows × 3 cols. Replace inputs with your real data when you run the converted output.
MorphSQL converted Snowflake → Python (pandas) (92% confidence).
For AI / ML practitioners
MorphSQL is a deterministic SQL→pandas / PySpark codegen tool (not a chat LLM). Use it when you have warehouse SQL for labels/features and want a Python frame for training or Spark jobs.
Typical path
- Convert feature SQL → pandas or PySpark
- Point
tables[...]at parquet /datasets/ warehouse extracts / Spark tables - Feed
resultinto sklearn, XGBoost, Transformers, or Spark ML
Why data scientists use this
Warehouse SQL often lives in BI tools. MorphSQL rewrites dialect quirks (NVL, ZEROIFNULL, dates) into pandas or PySpark you can run in Jupyter / Colab / Databricks.
Recommended workflow
- Paste SQL, load an example, or upload a
.sql/.zip - Choose Convert to (pandas / PySpark / Snowflake / BigQuery / dbt)
- Click Convert or Upload & Convert → Download
- Download the
.py/.sql/.zipto your machine and replace synthetic tables with real data
Output choices
| Convert to | Best for |
|---|---|
| Python (pandas) | Feature engineering, EDA, model training prep |
| Python (PySpark) | Large-scale Spark DataFrame transforms |
| Snowflake / BigQuery SQL | Keeping transforms in the warehouse |
| dbt project | Productionizing SQL into models |
More tab (Lab)
| Tool | Use when |
|---|---|
| Object assess | Need risk/complexity scoring for one SQL object |
| Repository workbench | Scanning a SQL repo (sample or zip) with lineage + validation |
| ML feature SQL | Migrating Vertica feature-engineering SQL → Snowflake/dbt |
| Copilot | Migration Q&A (keyword / optional HF LLM) |
| Behavior notes / Eval | Dialect quirks lookup and offline conversion scoring |
from morphsql.ai import pipeline
out = pipeline("sql-migration")(sql, source="snowflake", target="pandas")
# or target="pyspark"
# out["converted_sql"] → exec / save as features.py
Lab extras beyond day-to-day Convert: assess objects, scan a repo, migrate feature SQL, ask the copilot, or run offline eval.
Score complexity/risk and convert a single SQL object (pandas, PySpark, warehouse SQL, or dbt).
Scan the sample Vertica repo (or upload a .zip), then convert / validate / preview dbt + lineage.
Convert examples/ml_features/churn_feature_sql.sql (Vertica feature engineering) → Snowflake SQL or a dbt feature mart.
Ask migration questions. Uses keyword/HF fallback guidance (set HF_TOKEN for LLM replies).
Leaderboard
| Rank | Name | Pass rate | Token F1 | Exact | Pairs |
|---|---|---|---|---|---|
| 1 | smoke-test | 100.0% | 99.9% | 93.3% | 30 |