f*ck - Fields Combined with Columnar Keys
Universal Columnar Data Merging Tool
f*ck is a powerful Rust-based data merging engine that empowers users to combine, clean, and transform messy tabular data through an intuitive DSL and visual interface.
What is f*ck?
f*ck stands for "fields combined with columnar keys" - the core concept of merging data fields across multiple sources using columnar key relationships with intelligent merge policies.
Key Features
- π Smart Joins: Dynamic column mapping between different data sources
- π Aggregation Policies: Sum, Count, Average, Min, Max, FirstMatch
- π― Primary Key Logic: OR/AND logic for complex key relationships
- β‘ Lazy Evaluation: Powered by Polars for efficient processing
- π Incremental Computation: Salsa-based caching for performance
- π Multi-Modal: CLI, Daemon+RPC, and WASM support
- π Visual DSL: JSON-based query language
Quick Start
Installation
Basic Usage
- Prepare your data sources (CSV, TSV, XLSX, SQLite)
- Create a query plan (JSON DSL)
- Execute the merge
# Preview results
# Write to file
Example: Customer Order Analysis
Input Files
customers.csv
id,name,email
1,John Doe,john@example.com
2,Jane Smith,jane@example.com
3,Bob Johnson,bob@example.com
orders.csv
customer_id,order_total,product
1,99.99,Widget A
2,149.50,Widget B
1,25.00,Widget C
Query Plan (query.json)
Output
customer_id,customer_name,email,total_spent
1,John Doe,john@example.com,124.99
2,Jane Smith,jane@example.com,149.50
3,Bob Johnson,bob@example.com,0.0
Merge Policies
| Policy | Description | Use Case |
|---|---|---|
FirstMatch |
Take first non-null value | Contact info, names |
Sum |
Add all values | Order totals, quantities |
Count |
Count non-null entries | Number of transactions |
Average |
Mean of all values | Average order size |
Min |
Minimum value | Earliest date, lowest price |
Max |
Maximum value | Latest date, highest price |
CLI Options
Architecture
βββββββββββββββββββ βββββββββββββββββββ βββββββββββββββββββ
β Data Sources β β Query DSL β β Output β
β β β β β β
β β’ CSV/TSV βββββΆβ β’ Field Maps βββββΆβ β’ CSV/TSV β
β β’ XLSX β β β’ Join Logic β β β’ XLSX β
β β’ SQLite β β β’ Merge Policy β β β’ SQLite β
βββββββββββββββββββ βββββββββββββββββββ βββββββββββββββββββ
Core Components
- DSL Engine: JSON-based query planning and validation
- Data Reader: Multi-format input with Polars lazy evaluation
- Join Engine: Dynamic column mapping and transitive closure
- Aggregation Engine: Group-by operations with merge policies
- Output Writer: Multi-format export with streaming
Roadmap
Phase 1: Core Engine β
- Basic CSV join functionality
- DSL query planning
- Aggregation policies (sum, count, etc.)
- CLI interface
Phase 2: Advanced Features π§
- Salsa incremental computation
- WASM compilation support
- Transitive closure joins
- Type detection heuristics
Phase 3: UI & Integration π
- Web-based visual interface
- Real-time preview system
- Data lineage tracking
- Recipe sharing
Contributing
- Fork the repository
- Create a feature branch:
git checkout -b feature/amazing-feature - Commit changes:
git commit -m 'Add amazing feature' - Push to branch:
git push origin feature/amazing-feature - Open a Pull Request
Development
# Build and test
# Run with sample data
# Check WASM compatibility (currently limited)
License
This project is licensed under the MIT License - see the LICENSE file for details.
Why "f*ck"?
The name represents both the frustration of working with messy data and the satisfaction of finally getting it clean. f*ck is about taking control of your data and making it work for you.
"fck around and find out... how clean your data can be."*