wc और uniq
In this page:
wc [-l | -w | -c] file_name
sort file_name | uniq [-c]
Lines, Words, और Characters Count करना
wc line, word, और byte counts report करता है। -l, -w, या -c pass करना output को सिर्फ उस एक count तक restrict करता है, जो default three-column output parse करने से scripts में उपयोग करना आसान है।
< से redirect करना count के साथ filename print होने से बचाता है।उदाहरण: Counting Lines, Words, and Characters
#!/bin/bash
printf "one two\nthree\n" > sample.txt
echo "lines: $(wc -l < sample.txt)"
echo "words: $(wc -w < sample.txt)"
rm -f sample.txt
Login to try C/C++/Java/PHP code in the editor
Adjacent Duplicates हटाना
uniq identical consecutive lines के runs को एक में collapse कर देता है। यह file में non-adjacently scattered duplicates को REMOVE नहीं करता, इसलिए input को usually पहले sort किया जाता है duplicates को साथ लाने के लिए।
उदाहरण: Removing Adjacent Duplicates
#!/bin/bash
printf "a\na\nb\na\nb\nb\n" > raw.txt
echo "uniq alone (misses non-adjacent duplicates):"
uniq raw.txt
echo "sort then uniq (all duplicates collapsed):"
sort raw.txt | uniq
rm -f raw.txt
Login to try C/C++/Java/PHP code in the editor
Occurrences Count करना
uniq -c output की हर line को उस संख्या से prefix करता है जितनी बार वह consecutively appear हुई, items की एक list को एक frequency table में बदलते हुए।
उदाहरण: Counting Occurrences
#!/bin/bash
printf "apple\nbanana\napple\napple\nbanana\n" | sort | uniq -c
Login to try C/C++/Java/PHP code in the editor
सबसे Frequent Items ढूंढना
sort | uniq -c | sort -n को chain करना items को sort करता है, consecutive duplicates count करता है, फिर उस count से numerically sort करता है, एक ranked frequency list produce करते हुए, एक बहुत common one-liner।
उदाहरण: Finding the Most Frequent Items
#!/bin/bash
printf "a\nb\na\nc\na\nb\n" | sort | uniq -c | sort -n
Login to try C/C++/Java/PHP code in the editor
- यह भूल जाना कि
uniqसिर्फ ADJACENT duplicate lines हटाता है, file में कहीं भी के सभी duplicates नहीं; input को usually पहले sort करने की ज़रूरत होती है। wc -lउपयोग करना और यह भूल जाना कि यह newline characters count करता है, इसलिए एक trailing newline missing वाली file expected से एक line कम report कर सकती है।- यह न जानना कि
uniq -cहर line को prefix करता है कितनी बार वह occur हुई, जोsortके साथ एक बहुत common combination है।
wc -l,wc -w,wc -cक्रमशः lines, words, और bytes count करते हैं।uniqसिर्फ consecutive duplicate lines collapse करता है, इसलिए input typically पहलेsortसे pipe किया जाता है।uniq -cहर output line को prefix करता है एक count से कि वह कितनी बार repeat हुई।sort | uniq -c | sort -nfrequency counting के लिए एक classic combo है।
Chapter Quiz — Complete all 12 topics to unlock
0/12 topics done
Complete these topics first: