← Back to Bash Course | Chapter 10: Text Processing Tools | Lesson 9 of 12

wc और uniq

wc text में lines, words, या characters count करता है, और uniq repeated adjacent lines को remove या report करता है।
Syntax
bash
wc [-l | -w | -c] file_name
sort file_name | uniq [-c]

Lines, Words, और Characters Count करना

wc line, word, और byte counts report करता है। -l, -w, या -c pass करना output को सिर्फ उस एक count तक restrict करता है, जो default three-column output parse करने से scripts में उपयोग करना आसान है।

Note: File को argument के रूप में pass करने के बजाय < से redirect करना count के साथ filename print होने से बचाता है।

उदाहरण: Counting Lines, Words, and Characters

bash
#!/bin/bash
printf "one two\nthree\n" > sample.txt
echo "lines: $(wc -l < sample.txt)"
echo "words: $(wc -w < sample.txt)"
rm -f sample.txt

Adjacent Duplicates हटाना

uniq identical consecutive lines के runs को एक में collapse कर देता है। यह file में non-adjacently scattered duplicates को REMOVE नहीं करता, इसलिए input को usually पहले sort किया जाता है duplicates को साथ लाने के लिए।

उदाहरण: Removing Adjacent Duplicates

bash
#!/bin/bash
printf "a\na\nb\na\nb\nb\n" > raw.txt
echo "uniq alone (misses non-adjacent duplicates):"
uniq raw.txt
echo "sort then uniq (all duplicates collapsed):"
sort raw.txt | uniq
rm -f raw.txt

Occurrences Count करना

uniq -c output की हर line को उस संख्या से prefix करता है जितनी बार वह consecutively appear हुई, items की एक list को एक frequency table में बदलते हुए।

उदाहरण: Counting Occurrences

bash
#!/bin/bash
printf "apple\nbanana\napple\napple\nbanana\n" | sort | uniq -c

सबसे Frequent Items ढूंढना

sort | uniq -c | sort -n को chain करना items को sort करता है, consecutive duplicates count करता है, फिर उस count से numerically sort करता है, एक ranked frequency list produce करते हुए, एक बहुत common one-liner।

उदाहरण: Finding the Most Frequent Items

bash
#!/bin/bash
printf "a\nb\na\nc\na\nb\n" | sort | uniq -c | sort -n
Related Topics
{# common_mistakes/chapter_summary/browser_support: on Hindi pages the view already swaps in the hi_ translation fields (or blanks these out if untranslated), so this renders correctly for both languages without a lang_code check here. #}
आम गलतियां
  1. यह भूल जाना कि uniq सिर्फ ADJACENT duplicate lines हटाता है, file में कहीं भी के सभी duplicates नहीं; input को usually पहले sort करने की ज़रूरत होती है।
  2. wc -l उपयोग करना और यह भूल जाना कि यह newline characters count करता है, इसलिए एक trailing newline missing वाली file expected से एक line कम report कर सकती है।
  3. यह न जानना कि uniq -c हर line को prefix करता है कितनी बार वह occur हुई, जो sort के साथ एक बहुत common combination है।
चैप्टर सारांश
  • wc -l, wc -w, wc -c क्रमशः lines, words, और bytes count करते हैं।
  • uniq सिर्फ consecutive duplicate lines collapse करता है, इसलिए input typically पहले sort से pipe किया जाता है।
  • uniq -c हर output line को prefix करता है एक count से कि वह कितनी बार repeat हुई।
  • sort | uniq -c | sort -n frequency counting के लिए एक classic combo है।
🔒

Chapter Quiz — Complete all 12 topics to unlock

0/12 topics done

Complete these topics first:

Login to run this code

C/C++/Java/PHP execution requires a free account. Your code is saved — you'll land right back in the editor after logging in.